
BYOK AI Workspace for Teams: Compare Models | BounceGrip
在安全的共享团队工作区中比较和使用多个LLM模型,使用自己的API密钥。
@uncoolavatar · X
完整作品展
技术栈
60 projects

在安全的共享团队工作区中比较和使用多个LLM模型,使用自己的API密钥。
@uncoolavatar · X

查看LLM模型在10个基准问题上的评分和排名。
fristovic · HN
She watched me look at model rankings and asked what do the numbers mean... I literally had no good way of explaining it to her so I just came up with something that is approximately in the same ballpark as some of the benchmarks out there lol

查看和对比主流AI模型的公众意见和基准评分。
u/TasteMysterious5285 · Reddit
I built AI Census, a live field bulletin for how people are actually talking about AI models I’ve been building AI Census, a public “field bulletin” for how people are talking about current AI models. I kept running into the same problem: benchmark tables tell me how a model performs on a test, but not whether people are actually finding it useful, frustrating, reliable, etc. So I built a rolling view from public technical conversations across Reddit, Hacker News, Bluesky, GitHub, and Huggi

上传供应商合同进行 AI 并排比较和风险检测。
@nbOlveira · X
A vendor comparison software

Compare up to 10 vehicles, configure a build, and dig into performance, EV charging, fuel and maintenance costs with a single automotive data platform.
@MohammadAa7w81 · X
Check out what I just built with Lovable!

多模型事实检验API,在部署前验证AI输出的准确性。
kostaj · Product Hunt
Lenz Independent, multi-model fact-checking API for AI workflows

XTokenChecker是一个AI模型目录,用于对比不同AI网关上的覆盖率、定价和延迟。
zizheruan · HN
XTokenChecker – Verifies model identities of your AI gateway

AI fashion photography platform for e-commerce: model swap, flat-lay to on-model, garment recolor, and AI packshots, with pixel-perfect garment preservation.
@8DavideRighini8 · X

用0-100 AGI分数对标前沿AI模型的基准性能。
baraklaniado · HN
I audited my AI leaderboard scale – every score dropped 6-15 points

检测LLM API是否被降智或偷换模型,一键跑6项探针得出结果
cocodot LLM 降智检测 — 免费的 LLM API「降智/偷换模型」在线检测:填入任意 OpenAI 兼容端点的 base_url 和临时 API Key,跑 6 项探针(模型声明、动态题、能力完整性等)生成分项报告;Key 仅用于当次检测、不落库不留存,检测方法[开源](https://github.com/cocodot2026/cocodot-llmprobe)

在OpenVibeEval中对比不同AI模型生成前端代码和可访问性评分。
u/12qwww · Reddit
I built a live benchmark to see which AI actually writes the best frontend code Hey everyone! I built OpenVibeEval because I was tired of "vibe-checking" AI-generated frontend code. I wanted to know which model actually produces the most accessible and clean React/Tailwind output. What I built: •A leaderboard of 24 models (Claude, GPT, DeepSeek, etc.) ranked by axe-core accessibility scores. •A Harness Comparator to show how different system prompts change the same model's output. •

对比450+个LLM API定价方案,计算实际月度成本,支持缓存和批处理定价。
u/Greywolff06 · Reddit
I built LLMPrice — a free calculator for comparing LLM API costs across 450+ pricing routes I kept running into the same problem when comparing LLM APIs: the headline token price doesn't always tell you what your actual workload will cost. Caching, batch pricing, reasoning tokens, retries, different endpoints, and OpenRouter routes can change the result quite a bit. So I built LLMPrice.com. You enter your workload once — requests, input/output tokens, caching, retries, etc. — and it com