
System 2 Arena - Objective AI Strategy Benchmarks
与AI模型进行策略游戏,查看大语言模型在排行榜上的排名。
masterchef2209 · HN
I created a platform to check which AI models is the best gamer
完整作品展
技术栈
25 projects

与AI模型进行策略游戏,查看大语言模型在排行榜上的排名。
masterchef2209 · HN
I created a platform to check which AI models is the best gamer

为LLM输出提供token级引文API,通过注意力分析验证。
apoorvumang · HN
TokenPath – token-level citations for LLM output, read from attention

通过AI驱动的模拟面试场景练习系统设计、DSA和编程面试。
@ShivanshAg50455 · X

对版本控制系统和编码代理进行性能基准测试。
videlov · HN
I was interested in answering this question so I built a benchmark comparing git, jj and gitbutler in agentic context https://vcbench.dev/ Disclaimer - I am a co-founder of GitButler

检查 RAG 块并可视化 AI 代理工作流、内存架构和执行轨迹。
@Higgs0110 · X

在智能代理笔记本中构建、运行和评估机器学习工作流。
eldar_hsnv · HN
Show HN: AI Notebook for Data Science – Kind of Like Cursor but for Jupyter

用独立验证的数据对比和评估 SaaS API,辅助自建或购买决策。
fenilsuchak · HN
OpenBenchmarks – Helping agents discover and pick the right SaaS APIs

通过InferAll统一API访问207+个AI模型
TaylorM492 · HN
InferAll – One API for OpenAI, Anthropic, Google, Nvidia Nim

Agent原生的TypeScript框架,在托管GPU上训练和部署定制模型。
@soleil_colza_ · X

探索Y Combinator公司,获取你的创业想法与加速器契合度的AI分析。
sirily11 · HN
Application Signal – AI That Evaluates Your YC Startup Idea

监控AI代理的LLM调用、API和基础设施性能。
kirankgollu · HN
Oodle.ai – $10 per million agent traces
