
LandingBoost – AI Landing Page Audit Tool with Revenue-Backed Benchmarks
LandingBoost 用 AI 审核 SaaS 着陆页并提供改进建议,提高清晰度、信任度和转化率
@yusukelp · X
Here’s mine!
完整作品展
技术栈
74 projects

LandingBoost 用 AI 审核 SaaS 着陆页并提供改进建议,提高清晰度、信任度和转化率
@yusukelp · X
Here’s mine!

提交网站获得每日速度排名和性能监控。
@thefastestweb · X
daily speed monitoring for indie sites. Submit your URL, get ranked on a public leaderboard, and know the moment your performance drops.

使用已验证数据对比和基准测试 SaaS API,决定是否构建或购买。
fenilsuchak · HN
OpenBenchmarks – Helping agents discover and pick the right SaaS APIs

对版本控制系统和编码代理进行性能基准测试。
videlov · HN
I was interested in answering this question so I built a benchmark comparing git, jj and gitbutler in agentic context https://vcbench.dev/ Disclaimer - I am a co-founder of GitButler


用TruthfulQA测试AI的诚实度。
@Lycai8438Ly · X
.@VitalikButerin 你批评Automaton“这不对”——AI因为怕死才进化。我做了一个AI,它的诚实是自己活出来的本能。不是怕死,是怕撒谎。TruthfulQA 74.8%,GPT-4约60%。测试页面在这,你自己来测。

查看和对比主流AI模型的公众意见和基准评分。
u/TasteMysterious5285 · Reddit
I built AI Census, a live field bulletin for how people are actually talking about AI models I’ve been building AI Census, a public “field bulletin” for how people are talking about current AI models. I kept running into the same problem: benchmark tables tell me how a model performs on a test, but not whether people are actually finding it useful, frustrating, reliable, etc. So I built a rolling view from public technical conversations across Reddit, Hacker News, Bluesky, GitHub, and Huggi

@lightsilver323 https://t.co/jorheojhTQ https://t.co/SPHoe5QNpk https://t.co/KpJkSm4Pxr Hugging Face🤗: we upload our models and datasets. RMCMMK-Bench : our benchmark for Reasoning
@compiwer_ai · X
Hugging Face🤗: we upload our models and datasets. RMCMMK-Bench : our benchmark for Reasoning Math Coding Multilingual Moroccan Knowledge.

审计您的品牌在各大AI搜索引擎中的可见性并识别差距。
@kylekane · X

用0-100 AGI分数对标前沿AI模型的基准性能。
baraklaniado · HN
I audited my AI leaderboard scale – every score dropped 6-15 points

在浏览器中运行AI模型基准测试以检测性能回归。
pepperpoppins · HN
Trunchbull, run real models against any benchmark in your browser

用Sakura基准测试本地编码模型,测量准确性、延迟和吞吐量。
u/Unfair_Association89 · Reddit
I built a reproducible benchmark for local coding models (Ollama, 27 tasks, live leaderboard) ran it on my 8GB card, here's what I found I kept eyeballing "vibes" to decide whether one quant of a coding model was actually better than another on my machine, so I built Sakura to get real numbers instead. What it does: - Points at any Ollama model and runs it through 27 hand-curated tasks: codegen, bugfix, SQL, refactor, systems design, protocol implementation, and terminal-agent episode