
OmniCorp · 诚实测试
用TruthfulQA测试AI的诚实度。
@Lycai8438Ly · X
.@VitalikButerin 你批评Automaton“这不对”——AI因为怕死才进化。我做了一个AI,它的诚实是自己活出来的本能。不是怕死,是怕撒谎。TruthfulQA 74.8%,GPT-4约60%。测试页面在这,你自己来测。
完整作品展
技术栈
61 projects

用TruthfulQA测试AI的诚实度。
@Lycai8438Ly · X
.@VitalikButerin 你批评Automaton“这不对”——AI因为怕死才进化。我做了一个AI,它的诚实是自己活出来的本能。不是怕死,是怕撒谎。TruthfulQA 74.8%,GPT-4约60%。测试页面在这,你自己来测。

The consent-only public security scoreboard. Rank is earned by passive security signals, never by money or likes.
@Noumenon_ai · X
OutClean isn't about the money. It isn't about the rank. It's about whether the site holding your account, your product, your users, and your clients is actually secure. That's the only thing worth ranking.

简历检查和AI职业测试工具,评分您的工作安全指数及自动化风险。
@_b_meet · X

评估AI应用是否可交付、脆弱或实验项目的免费评分卡。
@eric810829 · X
I made a free Vibe-Coded App Reality Score. It checks the boring launch proof: - payment/access state - private data boundaries - AI output evidence - support recovery - one real user flow 0-100 score:

El ranking publicitario que arranca en US$1.00 y sube 25% para tomar el #1. Sin sorteos y con ranking transparente.
@astridmusexl · X

Public token leaderboard ranked by bid weight. Highest score wins. Score decays over time.
@Ranktokenlol · X
Built a live board where rank is a dollar amount that melts ~2.5% an hour.

查看和对比主流AI模型的公众意见和基准评分。
u/TasteMysterious5285 · Reddit
I built AI Census, a live field bulletin for how people are actually talking about AI models I’ve been building AI Census, a public “field bulletin” for how people are talking about current AI models. I kept running into the same problem: benchmark tables tell me how a model performs on a test, but not whether people are actually finding it useful, frustrating, reliable, etc. So I built a rolling view from public technical conversations across Reddit, Hacker News, Bluesky, GitHub, and Huggi

投票参与Groundbeat的实时民意调查,按答题者分类查看结果。
dgerken · HN
I rebuilt my 2000s polling site with an AI running the editorial desk

浏览并安装1,800+个MCP服务器,按信任度和风险等级排名。
@Shidesheng0218 · X
🔥 Just launched: MCP Market 1,600+ MCP servers ranked by Trust Score, Risk Level, and install success data. For Cursor, Claude, VS Code, and Codex. Copy config. Know it works. No more guessing. #MCP #AITools #Cursor #Codex #BuildInPublic

AI测试工具,为你的颜值评分并提供个性化的美妆建议。
Free AI Beauty Test — AI 颜值测试,提供给你个性化的建议

See where money is at stake before you spend. Free one-home Health Score, or landlord pricing: 2–7 fixed £350 Decision Review; 8+ five-property sample; quoted pilot. Indicative ran
@EranHertz · X
The path towards energy independence starts with knowing what is going on, what’s known and what’s missing.

用0-100 AGI分数对标前沿AI模型的基准性能。
baraklaniado · HN
I audited my AI leaderboard scale – every score dropped 6-15 points