
Axon — the quality & FinOps layer for your AI agents
用 LLM 评估 AI agent 对话质量,提供评分卡和成本分析。
@tech_maju · X
完整作品展
技术栈
68 projects

用 LLM 评估 AI agent 对话质量,提供评分卡和成本分析。
@tech_maju · X

在这个10轮游戏中识别哪些文本由AI生成,对比真实历史散文。
sanj001 · HN
Can you tell Wodehouse from a model imitating Wodehouse?

设定目标和标准,由AI小组评估,裁判用引用来源做出判决。
nadermx · HN
Referee.Chat - Set the goal. An AI panel works. Referee clears it done

聊天对比 AI 模型,投票参与排行榜排名。
u/Rabus · Reddit
I got TestingModels too overcomplicated over the month it is running: looking for some feedback how to make it more useful and simpler I run a benchmark like arena.ai , but with pre-generated prompts. So far, nearly 6k people came in and like 30k comparisons has been made - which means the thing is genuinely useful for people to compare the models. The problem is the more features i started adding the more overblown and complicated UI became - like old internet explorer tab bars Old: ht

根据情绪追踪和行为分析为交易生成AI决策评分。
@eialgos · X
AI/ML builder here 👋 I'm building EI ALGOS, a Decision Intelligence platform that combines machine learning, behavioral analytics, options analysis, and technical evaluation to help traders make higher-quality decisions. Visit us at Happy to connect with other founders and builders.

用0-100 AGI分数对标前沿AI模型的基准性能。
baraklaniado · HN
I audited my AI leaderboard scale – every score dropped 6-15 points

用AI将目标拆分为每日结构化任务,通过积分和连续记录追踪进度。
u/AdConnect5584 · Reddit
Check out Goalforge https://goalforge.me/ Hi, recently build this project and it's basically goal tracker with integrated ai that helps turn general goals into smart structured ones. Also I included chatbot for those who wanna have a chat and maybe dive deeper and structure the goal more personally. Have fun, and let me know what breaks there submitted by /u/AdConnect5584 to r/SideProject [link] [comments]

Persistent memory API for AI agents. Hybrid scoring (semantic + recency + importance) that works with LangGraph, AutoGen, and CrewAI. Free to start.
@sunvic567 · X

用自动可读性评分检查 AI 代理对您网站的理解程度。
tommy2970 · HN
Lagotto Meter – how well do AI agents understand your website?

用AI根据具体性、证据质量和风险清晰度为创业想法评分。
@nellaiorgs · X
Building NELL Labs -- an AI-native startup validation platform. Instead of generic LLM answers, it scores your idea across specificity, evidence quality, risk clarity, and next-step usefulness, so you know if it's actually validated or just sounds good.

免费AI评估,找出孩子的学习优势和不足,推荐个性化Khan Academy课程。
@edsull · X
My vibe coding team of agents set up a series of AI assessments for each subject and each grade level. Check it out!

ScoreClash - Predict scores, use augments, and climb the leaderboard!
@ScoreClash · X