
AI Benchmark Leaderboards & Model Evals | BenchmarkList
Compare and evaluate AI models across coding, reasoning, agents, and other benchmarks.
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities
The full gallery
Tech stack
60 projects

Compare and evaluate AI models across coding, reasoning, agents, and other benchmarks.
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

Study Kubernetes, AI, Linux, and DevOps with spaced repetition flashcards.
Agoreddah · HN
I built Gnoseed – free flahscards to learn K8s, DevOps, AI and more

Assess your B2B AI or SaaS startup's biggest bottleneck blocking your next milestone.
@FounderUnstuck · X
If you’re building a startup and want to identify your biggest bottleneck, try the free assessment: 🔗 We’d love to hear if the results match your experience.

Compare screen recording software by features, pricing, and performance benchmarks.
@y276161014 · X
有人做了一个屏幕录制软件对比评分网站。 太需要了!

Benchmark AI honesty with TruthfulQA test questions.
@Lycai8438Ly · X
.@VitalikButerin 你批评Automaton“这不对”——AI因为怕死才进化。我做了一个AI,它的诚实是自己活出来的本能。不是怕死,是怕撒谎。TruthfulQA 74.8%,GPT-4约60%。测试页面在这,你自己来测。

Submit your website for daily speed benchmarking and ranking on a public leaderboard.
@thefastestweb · X
daily speed monitoring for indie sites. Submit your URL, get ranked on a public leaderboard, and know the moment your performance drops.

Portable memory that persists across Claude, ChatGPT, Cursor, and other AI assistants.
u/OrganicArgument2092 · Reddit
I shipped Lodekeep: portable memory for AI agents that follows you across Claude, Cursor, and ChatGPT Every new AI chat starts from zero. Claude, Cursor, ChatGPT all forget my stack, my decisions, the gotchas I already solved, so I kept re-pasting the same context every single session. I got sick of it and built Lodekeep. You capture a decision, preference, or lesson once, and it's recallable in every future session across every MCP client (Claude web + desktop, Claude Code, Cursor, Gemini

WitBench.com: AI sense of humor benchmark I created witbench.com benchmark because everyone's measuring math and code performance, but personally, I like laughing. TL;DR: Gemi
u/tziki · Reddit
WitBench.com: AI sense of humor benchmark I created witbench.com benchmark because everyone's measuring math and code performance, but personally, I like laughing. TL;DR: Gemini funny, Grok unfunny, but do check out the full list, I spent real money on actual impartial raters. submitted by /u/tziki to r/SideProject [link] [comments]

Write once, export to notebooks, slides, PDFs, PPTX, DOCX, or canvas.
@mr_wickedhacks · X
Write the doc once → get a notebook, slide deck, canvas, or resume from the same source. No rewriting for every format.

File-first memory runtime for AI agents with dashboard and HTTPS API access.
@memofsdev · X

Browser-based tools for formatting, encoding, testing, and converting data.
4seasons · V2EX
devfun.org - 开发者(网络)工具集合 ## 访问地址 https://devfun.org ## 主要功能 - 本机 IPv4/IPv6 网络检查及 IP 信息查询 - 站点可访问性检查(支持全球区域) - 代理检查及代理全球出口检查 - 各类编解码、格式测试、文本格式化、校验密钥生成、文本处理、转换工具 欢迎大家使用,提供建设性的建议和意见!

Store shared memories once and reuse them across AI systems and your team.
Repeater22746 · HN
ContextVault – Shared memory layer for your AI and your team