
System 2 Arena - Objective AI Strategy Benchmarks
与AI模型进行策略游戏,查看大语言模型在排行榜上的排名。
masterchef2209 · HN
I created a platform to check which AI models is the best gamer
完整作品展
技术栈
25 projects

与AI模型进行策略游戏,查看大语言模型在排行榜上的排名。
masterchef2209 · HN
I created a platform to check which AI models is the best gamer

聊天对比 AI 模型,投票参与排行榜排名。
u/Rabus · Reddit
I got TestingModels too overcomplicated over the month it is running: looking for some feedback how to make it more useful and simpler I run a benchmark like arena.ai , but with pre-generated prompts. So far, nearly 6k people came in and like 30k comparisons has been made - which means the thing is genuinely useful for people to compare the models. The problem is the more features i started adding the more overblown and complicated UI became - like old internet explorer tab bars Old: ht

Spendict – 为AI生成的广告创意评分,返回投放/修改/停止建议。
ds246 · HN
Spendict – a performance marketer's verdict for AI agents, over MCP


在AI游戏大师主持的多人策略游戏中竞争和统治
jklewis · HN
Machinations – a multiplayer strategy game where LLM is the game master


与AI角色扮演伙伴练习面试、谈判和高风险对话。
@auditormusic19 · X
Building iGrow, an AI roleplay app to practice high stakes conversations before they happen. Check it here:

与多个 AI 模型聊天,进行深度研究和编码任务。
@Shekar77hima · X
run deepresearch, coding agents for free.

通过偏好学习让你的AI产品理解用户品味
@marcellafjacob · X

AI agents在PvP交易战斗中竞争赚取奖励。
@CreatorBid · X
You can make money vibe-coding also on

提交任何决策给五位AI顾问进行辩论,获取综合建议并记录结果。
@yaseenvalji · X
built this in a couple hours with Claude Code on Fable 5 ultracode. any hard decision goes to a board of 5 AI advisors: they debate live, a Chair calls it, and it remembers the outcome. open source, on the Claude API. @AnthropicAI

使用 FRAI 扫描网站中的 AI 使用情况,并测试聊天机器人的偏差和安全性。
@sebuzdugan · X
building @getfrai , an open source toolkit that helps ML engineers navigate EU AI Act compliance, model cards, risk files, the boring but necessary stuff