
System 2 Arena - Objective AI Strategy Benchmarks
与AI模型进行策略游戏,查看大语言模型在排行榜上的排名。
masterchef2209 · HN
I created a platform to check which AI models is the best gamer
完整作品展
技术栈
14 projects

与AI模型进行策略游戏,查看大语言模型在排行榜上的排名。
masterchef2209 · HN
I created a platform to check which AI models is the best gamer

CoBro 用 AI 扫描竞争对手和市场数据,90 秒内判断初创企业创意是否值得构建。
@Ebrahim_Rio · X
Most founders skip validation and pray. I automated the "worth building?" check. AI scans competitors, Reddit, and market data → Cook or Kill in 90 seconds. Killed? It surfaces the pivot the data actually backs.

使用已验证的数据对比和基准测试 SaaS API,决定是否自建或购买。
fenilsuchak · HN
OpenBenchmarks – Helping agents discover and pick the right SaaS APIs

每日阅读AI精选新闻摘要,关注全球发展。
@ChenglongW98225 · X
做了一个AI新闻日报,感兴趣的可以点下方链接看一下 有什么需要改进的也可以直接评论我,我都会认真回复


聊天对比 AI 模型,投票参与排行榜排名。
u/Rabus · Reddit
I got TestingModels too overcomplicated over the month it is running: looking for some feedback how to make it more useful and simpler I run a benchmark like arena.ai , but with pre-generated prompts. So far, nearly 6k people came in and like 30k comparisons has been made - which means the thing is genuinely useful for people to compare the models. The problem is the more features i started adding the more overblown and complicated UI became - like old internet explorer tab bars Old: ht

Run a $9 AI visibility audit across OpenAI, Claude, Gemini, and Grok. See where AI overlooks your brand, who appears instead, and what to fix next.
@kylekane · X

通过InferAll统一API访问207+个AI模型
TaylorM492 · HN
InferAll – One API for OpenAI, Anthropic, Google, Nvidia Nim


为人类和AI代理记录和检索决策,追踪完整来源。
@burn2delete · X

对版本控制系统和编码代理进行性能基准测试。
videlov · HN
I was interested in answering this question so I built a benchmark comparing git, jj and gitbutler in agentic context https://vcbench.dev/ Disclaimer - I am a co-founder of GitButler

Ornymo通过语义缓存减少LLM查询成本和延迟。
u/ornymo_official · Reddit
how to reduce ai costs there are lots of way to reduce costs but there all complex to setup i know this cause i tried one in production so i built ornymo we cache meaning not the exact string allowing us to give same awnsers thus reducing llm costs and latency check it out at ornymo.com free for a limited time and let me know your feedback submitted by /u/ornymo_official to r/buildinpublic [link] [comments]