
Telemetry — Observability for AI and LLM apps
监控AI应用中的模型调用、代理步骤和检索,追踪令牌、成本和延迟。
ephraimduncan · HN
Observability for Coding Agents and LLM Applications
完整作品展
技术栈
60 projects

监控AI应用中的模型调用、代理步骤和检索,追踪令牌、成本和延迟。
ephraimduncan · HN
Observability for Coding Agents and LLM Applications

分析新闻文章,验证结论是否得到证据支持。
@BiaoBuilds · X
我做了一个帮助读者拆解新闻论证结构的工具:LedeLens。 它不做事实核查,也不判断政治倾向,只回答一个更小的问题:文章的结论,能否由它自己提供的证据支持? 在线体验: 感兴趣可以看看~

用对抗测试检查LLM端点安全,获取OWASP审计报告。
@aryaan_sheth · X
- LLM security for small teams

检测LLM API是否被降智或偷换模型,一键跑6项探针得出结果
cocodot LLM 降智检测 — 免费的 LLM API「降智/偷换模型」在线检测:填入任意 OpenAI 兼容端点的 base_url 和临时 API Key,跑 6 项探针(模型声明、动态题、能力完整性等)生成分项报告;Key 仅用于当次检测、不落库不留存,检测方法[开源](https://github.com/cocodot2026/cocodot-llmprobe)

为LLM输出提供token级引文API,通过注意力分析验证。
apoorvumang · HN
TokenPath – token-level citations for LLM output, read from attention

多智能体LLM系统的可视化编辑器,支持本地推理。
sascha10000 · HN
Multi-agent LLM editor with local inference via WebSockets

在浏览器中运行AI模型基准测试以检测性能回归。
pepperpoppins · HN
Trunchbull, run real models against any benchmark in your browser

FlexInference: 通过多个提供商路由LLM API请求,降低成本和延迟。
Aperswal · HN
Made a Free LLM Router

用多个模型实时审计AI回应以判断其可靠性。
u/inc_23 · Reddit
Hey, I created a tool that catches when your LLM is confidently wrong, in production, in real time — looking for beta testers. Your bot sounds sure of itself even when it's wrong, and you usually only find out when a customer complains. Auscope audits every LLM response in the background: 3 models from 3 different providers independently check it, a 4th "chairman" model resolves disagreements, and you get one verdict — verified, uncertain, or unreliable. Runs async, doesn't slow your respon

通过 RavenGate 网关路由 LLM API 流量,追踪成本、分析延迟、隐蔽 PII。
charltonraven · HN
RavenGate – LLM gateway that redacts PII across SSE chunk boundaries


观看两个随机LLM在物理竞技场中剑战,盲投谁表现得更聪明。
u/Time-Shelter-35 · Reddit
I built a site where two LLMs sword-fight in real physics and you blind-vote who's smarter Two months ago I thought: what if the AI benchmark was just… watching them fight. So: https://stickblade-arena.vercel.app Two random LLMs get stickman bodies in a pymunk physics arena They each turn output JSON moves (swing, block, dash, shoot bow, throw flail…) Ragdolls, momentum, weapon collisions, the whole bit You watch the replay without knowing which model is which and vote who f