
AI Benchmark Leaderboards & Model Evals | BenchmarkList
对比和评估 AI 模型在编码、推理、代理和其他基准测试中的表现。
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities
完整作品展
技术栈
19 projects

对比和评估 AI 模型在编码、推理、代理和其他基准测试中的表现。
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

用 AI 创建产品路线图、管理任务并收集功能反馈。
@jimmy_harika · X
TLDR: Notion shipped my exact app that I have using in my daily workflow from last 2 years. Try here: It got a mcp that wires to your claude code and codex. Git integration is almost complete and will ship in coming days


隐私优先的网站分析平台,配备AI助手进行SEO内容研究和发布。
@VertCodeEU · X

用AI分析敏感数据,通过客户端加密实现端到端保护。
@JackiePeters · X

用自然语言问题查询电子表格和数据集,生成即时答案、报告和仪表板。
u/maybeImakemoney · Reddit
I built the thing. Now I am not sure the base use case is one people will pay for. Founder here. This started as a side learning project to see whether an LLM could answer questions about Excel data, back when they could not do it well. I built the first version on n8n, with workflows that ingested files, generated metadata with an LLM, and answered questions against the converted data plus that metadata. Then I started using it for my own analysis and report generation, saw that the time sav

使用 Kitbase 追踪用户事件并分析产品行为。
@kitbasedev · X
We just launched, check us out at

对比 AI 模型在编码任务上的表现,支持成本追踪和 ELO 排名。
@intheworldofai · X
On the World of AI Bench (vibe-coding composite): Claude Fable 5 → 85.2 GPT-5.6-sol → 82.4 kimi-k3 → 81.5 Moonshot’s K3 just walked in and claimed bronze on one of the toughest coding-focused leaderboards out there.

用ProductBlaze诊断产品市场契合度,发现增长瓶颈。
@teslaptimus · X

创建并启动 AI 服务,集成支付和分析功能。
EasyLaunch — 能力产品化平台,把专家方法论和 AI Skill 一键上线成带登录、收款与数据统计的可售服务

AI驱动虚拟数据室,帮助交易团队将散乱文档转化为引导性体验。
@Puneeeeeeet · X
We help brands go viral on X with organic video campaigns that people actually want to watch. wanna try for

上传 CAS PDF 获取 AI 驱动的基金组合分析和配置洞见。
@iASHeeesh · X