
AI Benchmark Leaderboards & Model Evals | BenchmarkList
对比和评估 AI 模型在编码、推理、代理和其他基准测试中的表现。
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities
完整作品展
技术栈
60 projects

对比和评估 AI 模型在编码、推理、代理和其他基准测试中的表现。
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

Ad-free Vedic astrology engine on the Swiss Ephemeris. Compute your sidereal chart, dasha timing and planetary strengths — then query it with AI.
@saketposwal · X

Match with investors that fit your startup, then improve your deck with investor-style feedback and readiness scoring. Start free and upgrade anytime.
@Evalyze_ai · X
Ran Shack again through to find the right pre-seed investors for consumer AI and automated selling. pulled 3 europe focused investors who actually back early consumer tech and marketplaces 1- @specht_p at Creandum. general partner who heavily backs early stage consumer platforms, marketplaces, and software across europe 2- @sophiabendz at Cherry Ventures. partner and former spotify exec who specifically targets early consumer tech and consumer AI apps 3- @chrija at Point Nine. managing partner who is a top early stage investor in marketplaces and scalable software drop your deck on the site to grab the rest of the list

使用自动化承保和敏感性分析来分析酒店投资。
u/Additional-Study2600 · Reddit
Yaay!!! I finally got some subscribers to my platform after 6 weeks… any growth tips? Hey yall I’m super excited because after months of vibe coding and sleepless nights learning about repos, PRs, branches and commits to main lol I finally have a product I’m proud of and my first real revenue! The platform is called Underwrote.AI. it’s B2B and it’s somewhat niche… it’s underwriting software for hotel acquisitions (think institutional-grade Excel models, generated from a guided workflow). O

检查 AI 推理踪迹,评估模型真实性。
malik_dixon1 · Product Hunt
TraceLogicAI: AI Architecture Evaluation Compare AI architectures with evidence, not guesswork

Whetstone 将 AI 候选方案与基准对比,否决回归并返回可审计的决策。
@JustinGarr90748 · X
We're building Cyberelf labs because a better score doesn't mean a better model.

查看LLM模型在10个基准问题上的评分和排名。
fristovic · HN
She watched me look at model rankings and asked what do the numbers mean... I literally had no good way of explaining it to her so I just came up with something that is approximately in the same ballpark as some of the benchmarks out there lol

The job board for AI developers — RAG, agents, evals, inference. First-party roles pulled nightly from company career sites.
alastairr · HN
Job board and MCP server for AI developers

AI fashion photography platform for e-commerce: model swap, flat-lay to on-model, garment recolor, and AI packshots, with pixel-perfect garment preservation.
@8DavideRighini8 · X

检测API中转站输出是否与官方100%一致
@nodeloc_cc · X
🌈 7月,你好,MODELOC上线算力池。 MODELOC自上线以来,已检测2000余次,覆盖600+中转站,为众多AI用户提供的使用参考。 MODELOC近期进行了改版,上线了算力池及市场。 加入算力池 查看帖子: 用 MODELOC 便宜地调各家大模型:一次讲清它的价格体系

让AI模型通过3D动画展现香蕉植物的完整生命周期来比较性能。
fran-mora · HN
I gave 5 AI coding agents one prompt: grow a banana plant through its whole life in three.js: sprout, leaves, flower, fruit, rot, then pups that restart the loop. It's deceptively simple and yet very hard to get right from procedural code: you have to write working three.js and understand how the plant is actually built; how it hangs, ages and decays. Get the biology wrong and the code renders something weird. These are agents, not bare models (Claude Code and Codex for now). They can use tools, including playwright to check their work and improve it.

V2Fun:从图像或文本生成3D模型、角色和动画
zaczuo · Product Hunt
V2Fun Generate 3D character with 8K textures and AI motion capture