
AI Benchmark Leaderboards & Model Evals | BenchmarkList
对比和评估 AI 模型在编码、推理、代理和其他基准测试中的表现。
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities
完整作品展
技术栈
60 projects

对比和评估 AI 模型在编码、推理、代理和其他基准测试中的表现。
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities


为任何主题生成动画视觉解释来学习。
u/Top-Relationship8196 · Reddit
AI teacher that explains anything visually. Hi, I built this free tool called bestie. It's basically an AI teacher that explains anythingany topic visually. Instead of watching a 3hr Youtube to learn something specific or struggle through the text Chatgpt gives you, you can use this for free to learn anything fast. you can find it at bestie.chrestic.com you just have to type the topic and your visuals will arrive in less than a minute. Don't forget to try and leave Y

LLM驱动的新闻聚合器,实时呈现和更新热门故事。
tdubey · HN
DWS A LLM Generated, "Drudge Report" style news site

评估知识后生成个性化课程的AI生成器
u/graywolf724 · Reddit
167 users, 202 courses, 1,840 chapters. $0 in revenue, and that's on purpose. https://reddit.com/link/1v94i9o/video/epyh6f5t40gh1/player Built an AI course generator that asks you ~5 adaptive questions to figure out what you actually already know, then builds a course around the gap instead of a generic curriculum. No account needed to try it, and it's still completely free (for now). The numbers so far: 167 people have generated at least one course, 202 courses total, about 1,840 ch

在浏览器中测试小型语言模型(8M-13M 参数),离线可用。
u/Live_Confusion_3003 · Reddit
I trained an LLM that runs on an ESP32 and directly in the browser Link to try it out yourself is: topk.sh The models download their weights directly in the browser so it works offline. Keep in mind they are very small and inaccurate. (8M and 13M parameters) However, I am building 500M and 1B+ parameter local models for agent based coding and other purposes. I will be shipping hardware designed for these tasks which connect directly to you computer or other device.

在OpenVibeEval中对比不同AI模型生成前端代码和可访问性评分。
u/12qwww · Reddit
I built a live benchmark to see which AI actually writes the best frontend code Hey everyone! I built OpenVibeEval because I was tired of "vibe-checking" AI-generated frontend code. I wanted to know which model actually produces the most accessible and clean React/Tailwind output. What I built: •A leaderboard of 24 models (Claude, GPT, DeepSeek, etc.) ranked by axe-core accessibility scores. •A Harness Comparator to show how different system prompts change the same model's output. •

An LLM gateway for OpenAI, Anthropic, Google and Azure. Every request logged, priced to the token, and audited for waste you can actually recover.
@razdagan3 · X

多智能体LLM系统的可视化编辑器,支持本地推理。
sascha10000 · HN
Multi-agent LLM editor with local inference via WebSockets

Ornymo通过语义缓存减少LLM查询成本和延迟。
u/ornymo_official · Reddit
how to reduce ai costs there are lots of way to reduce costs but there all complex to setup i know this cause i tried one in production so i built ornymo we cache meaning not the exact string allowing us to give same awnsers thus reducing llm costs and latency check it out at ornymo.com free for a limited time and let me know your feedback submitted by /u/ornymo_official to r/buildinpublic [link] [comments]

通过 RavenGate 网关路由 LLM API 流量,追踪成本、分析延迟、隐蔽 PII。
charltonraven · HN
RavenGate – LLM gateway that redacts PII across SSE chunk boundaries

对比多个LLM API提供商的延迟和吞吐量性能。
@QAInsights · X