
知行录 · leaderboard.cn
在 leaderboard 上按官方基准对比 AI 大模型的性能排名
fcten · V2EX
做了一个大模型 leaderboard 网站 最近一个月 CodeX 疯狂送重置,token 根本用不完,顺手做点东西。 地址:[知行录]( https://leaderboard.cn/) 排行依据主要为模型官方基准测试成绩。非主观排名。 数据会持续更新。如果有点用,欢迎各位 v 友收藏~
完整作品展
技术栈
60 projects

在 leaderboard 上按官方基准对比 AI 大模型的性能排名
fcten · V2EX
做了一个大模型 leaderboard 网站 最近一个月 CodeX 疯狂送重置,token 根本用不完,顺手做点东西。 地址:[知行录]( https://leaderboard.cn/) 排行依据主要为模型官方基准测试成绩。非主观排名。 数据会持续更新。如果有点用,欢迎各位 v 友收藏~

上传手写考试答卷,获得即时AI评分和详细反馈。
Xaminix
AI Powered Answer Evaluation for CA/CS/CMA

向多个前沿大模型提问,获得经过同行评审的综合答案。
u/Puzzleheaded-Log-27 · Reddit
Building a multi-model AI deliberation tool taught me something about trust LLM Counsel isn't another wrapper around one model - it sends your question to a panel of frontier LLMs, has them peer-review each other anonymously, and an impartial "chairman" model returns one synthesized answer. Free to start, pay-as-you-go after, credits don't expire. What I've learned so far: people trust a synthesized answer a lot more once they can see that the models actually disagreed and how that disagree

比较AI模型在多个领域的基准评估成绩和排行榜。
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

完成12道测试题,检验你对数字收入方法的了解。
@XERECRAZY4hll · X
Check out what I just built with Lovable!

将笔记转化为竞争性答题对战,获得AI反馈和指导。
u/Due_Load6189 · Reddit
I built a study battle app and need honest feedback Hey everyone, I’m a student building StudyClash, an app that turns practice questions into competitive study battles. You can try a deck, answer questions, get a score, see what topics you’re weak in, and use an AI coach to understand mistakes. I’m not trying to sell anything right now — I just need honest feedback before I show it to more people. Main things I want feedback on: Is the app easy to understand? Does anything look b

查看你的Vouch Score,衡量ChatGPT、Claude等AI对你的产品的推荐频率。
@abhishek_gadhia · X

向Claude、GPT和Gemini提问,获得它们相互审核的共识答案。
@StevenJdotCom · X
One AI makes mistakes. Three catch each other's. AI Consensus runs your prompt through Claude, GPT and Gemini. They work it independently, then critique each other until they reach consensus — handing you an AI audited, combined answer.

基于证据分析比较候选人,生成招聘决策文件。
facundobon · HN
Verdict – AI hiring verdicts where every score cites the CV verbatim

实时多人竞答游戏,配有AI主持人、30+游戏模式和考试备考包。
@deebee230730 · X

创建实时测验,AI即时生成问题,与观众实时互动。
EZQuiz — AI 智能实时测验,让你快速创建、分享并与观众通过实时互动测验进行互动

完成几个句子测量你的clanker评分。
niklio · HN
You write 8 text completions and open models score how predictable each word was too them. Predictable => clanker. You can share results with your friends. The scoring checks every word you write against the model's logprobs. Right now I'm using Llama3.1, Deepseek v3 and Qwen3 to keep costs low. I tried to calibrate it so other models (chatgpt/claude) score 100% and interesting human responses score in the 10-30% range. Totally free, no signup