
Mentiss | AI Werewolf, Mafia & Board Game Benchmark for LLMs
与AI对手玩狼人或黑手党桌游,对语言模型进行基准测试。
@jeremy_d_w_90 · X
做了一年多,Mentiss AI 狼人杀终于上线了 🐺 现在可以拉上朋友,和 AI 同桌发言、推理、互骗。 来玩一局: #AI狼人杀 #MentissAI
完整作品展
技术栈
60 projects

与AI对手玩狼人或黑手党桌游,对语言模型进行基准测试。
@jeremy_d_w_90 · X
做了一年多,Mentiss AI 狼人杀终于上线了 🐺 现在可以拉上朋友,和 AI 同桌发言、推理、互骗。 来玩一局: #AI狼人杀 #MentissAI

用Sakura基准测试本地编码模型,测量准确性、延迟和吞吐量。
u/Unfair_Association89 · Reddit
I built a reproducible benchmark for local coding models (Ollama, 27 tasks, live leaderboard) ran it on my 8GB card, here's what I found I kept eyeballing "vibes" to decide whether one quant of a coding model was actually better than another on my machine, so I built Sakura to get real numbers instead. What it does: - Points at any Ollama model and runs it through 27 hand-curated tasks: codegen, bugfix, SQL, refactor, systems design, protocol implementation, and terminal-agent episode

在 leaderboard 上按官方基准对比 AI 大模型的性能排名
fcten · V2EX
做了一个大模型 leaderboard 网站 最近一个月 CodeX 疯狂送重置,token 根本用不完,顺手做点东西。 地址:[知行录]( https://leaderboard.cn/) 排行依据主要为模型官方基准测试成绩。非主观排名。 数据会持续更新。如果有点用,欢迎各位 v 友收藏~

对比 AI 模型在编码任务上的表现,支持成本追踪和 ELO 排名。
@intheworldofai · X
On the World of AI Bench (vibe-coding composite): Claude Fable 5 → 85.2 GPT-5.6-sol → 82.4 kimi-k3 → 81.5 Moonshot’s K3 just walked in and claimed bronze on one of the toughest coding-focused leaderboards out there.

用多个AI模型同时翻译并对比结果。
orion1 · V2EX
vibe conding 撸了两个工具。 都是自己经常再用的 1 ,翻译站,模仿的 deepl ,主要是翻译网站只给一个翻译,有的时候未必是自己想要的那个,这个能调用多个模型都翻译一遍。 https://translate.v8ce.com/ 2 ,音视频处理工具,不上传数据,只调用浏览器能力。有的时候想简单裁切一下视频,压缩或者加速,打开大型处理软件有点太重了。 https://video-tools.v8ce.com/

对比和评估 AI 模型在编码、推理、代理和其他基准测试中的表现。
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

通过统一的 API 接口访问多个 AI 语言模型。
u/DanTahirCode · Reddit
I built an open source coding agent with a personality - meet Klenny Code 🐾 Hey r/SideProject, my name is Dan Tahir, and I'm here to show off something I'm really proud of: Klenny Code, the open source coding agent with personality. A fully capable coding agent with memory and cross-project referencing, plus an assistant who can read your email, run scheduled tasks, pilot your browser, and be your corgi pal. Here's the pitch: bring your own OpenRouter API key, and Klenny wil

发现你的电脑能运行的本地 AI 模型,包含已验证的基准数据。
@Carl0sFelipe · X
Just shipped — a tool that helps you discover which local AI models actually run on your hardware, with community benchmarks, quantization support, and estimated speed. Building in public from here. #BuildingPublic #AIDevelopment #rust #benchmaks #aimodel

@lightsilver323 https://t.co/jorheojhTQ https://t.co/SPHoe5QNpk https://t.co/KpJkSm4Pxr Hugging Face🤗: we upload our models and datasets. RMCMMK-Bench : our benchmark for Reasoning
@compiwer_ai · X
Hugging Face🤗: we upload our models and datasets. RMCMMK-Bench : our benchmark for Reasoning Math Coding Multilingual Moroccan Knowledge.

查看和对比主流AI模型的公众意见和基准评分。
u/TasteMysterious5285 · Reddit
I built AI Census, a live field bulletin for how people are actually talking about AI models I’ve been building AI Census, a public “field bulletin” for how people are talking about current AI models. I kept running into the same problem: benchmark tables tell me how a model performs on a test, but not whether people are actually finding it useful, frustrating, reliable, etc. So I built a rolling view from public technical conversations across Reddit, Hacker News, Bluesky, GitHub, and Huggi

在浏览器中测试小型语言模型(8M-13M 参数),离线可用。
u/Live_Confusion_3003 · Reddit
I trained an LLM that runs on an ESP32 and directly in the browser Link to try it out yourself is: topk.sh The models download their weights directly in the browser so it works offline. Keep in mind they are very small and inaccurate. (8M and 13M parameters) However, I am building 500M and 1B+ parameter local models for agent based coding and other purposes. I will be shipping hardware designed for these tasks which connect directly to you computer or other device.

探索语言模型如何解释和处理文本的交互式平台。
belluxx · HN
Nanointerpret – LLM Interpretability Playground