
AI Benchmark Leaderboards & Model Evals | BenchmarkList
对比和评估 AI 模型在编码、推理、代理和其他基准测试中的表现。
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities
完整作品展
技术栈
60 projects

对比和评估 AI 模型在编码、推理、代理和其他基准测试中的表现。
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

免费AI评估,找出孩子的学习优势和不足,推荐个性化Khan Academy课程。
@edsull · X
My vibe coding team of agents set up a series of AI assessments for each subject and each grade level. Check it out!

AI产业链股票池,支持美股/A股映射及一键部署。
yaoleifly · GitHub
ai-stock-pool AI industry-chain stock pool with US/A-share mapping, active discovery, policy pressure, and one-click deployment.

使用 FRAI 扫描网站中的 AI 使用情况,并测试聊天机器人的偏差和安全性。
@sebuzdugan · X
building @getfrai , an open source toolkit that helps ML engineers navigate EU AI Act compliance, model cards, risk files, the boring but necessary stuff

追踪全球AI工具每日使用趋势、按国家查看搜索偏好和SDK下载。
pixipace · HN
Who is using AI? A daily tracker, 1,600+ city heatmap, DOI dataset

用AI驱动的投资工具分析股票,包括DCF模型和选股器。
@StephanInvests · X
Building a financial saas at the lowest price for people to put more towards their financial goals

AI平台,提问、生成图像、语音交互,辅助更好的决策。
@Akinzoooo · X

追踪 AI 行业关键变化与证据支持的趋势。
barretlee · GitHub
agent-pulse Evidence-backed AI industry intelligence — trends, source updates, daily data refreshes, and weekly decision briefs.

查看和对比主流AI模型的公众意见和基准评分。
u/TasteMysterious5285 · Reddit
I built AI Census, a live field bulletin for how people are actually talking about AI models I’ve been building AI Census, a public “field bulletin” for how people are talking about current AI models. I kept running into the same problem: benchmark tables tell me how a model performs on a test, but not whether people are actually finding it useful, frustrating, reliable, etc. So I built a rolling view from public technical conversations across Reddit, Hacker News, Bluesky, GitHub, and Huggi

自主发现高价值排名机会并自动执行网站优化的 AI 系统。
@hola_gmi · X
built with claude ai. —

使用客户痛点、竞争信号和市场数据发现并验证AI产品机会。
vinitk80555 · HN
AI Product Opportunity

20 buyer-style prompts × 5 AI engines = 100 observation points per audit. Score 0-100. Free 60-second scan. Crypto-native methodology, every weight published.
@citeOS_io · X
is all yours