
LLM Status — AI model deprecation tracker & CLI checker
Track AI models used in your apps and receive warnings before they're deprecated.
taylorgt · HN
Find every AI model your code calls and warn before it's retired
The full gallery
Tech stack
60 projects

Track AI models used in your apps and receive warnings before they're deprecated.
taylorgt · HN
Find every AI model your code calls and warn before it's retired

Compress prompts before LLM API calls to reduce token usage and costs.
@asgujjuasitgets · X

Semantic caching reduces LLM token costs and latency for AI queries.
u/ornymo_official · Reddit
how to reduce ai costs there are lots of way to reduce costs but there all complex to setup i know this cause i tried one in production so i built ornymo we cache meaning not the exact string allowing us to give same awnsers thus reducing llm costs and latency check it out at ornymo.com free for a limited time and let me know your feedback submitted by /u/ornymo_official to r/buildinpublic [link] [comments]

Monitor AI model calls, agent steps, and retrieval with token tracking, cost analysis, and latency metrics.
ephraimduncan · HN
Observability for Coding Agents and LLM Applications

Use one API to access and switch between LLM providers while optimizing inference costs.
justin2025 · Product Hunt
Auriko Trading desk for LLM calls

Monitor Anthropic and OpenAI API calls with per-request token usage, costs, and session trees.
u/Prestige_pvp · Reddit
Simple AI Token Profiler / Debugger We made a simple profiler to help optimize AI token spend. Generally speaking anytime you want to optimize your app, whether it's for memory or otherwise you typically start with a profiler. There are a ton of MiTM Gateways but there aren't many true profilers, so I thought I'd make one. https://profiler.getrekon.com/ Let me know what you think :) submitted by /u/Prestige_pvp to r/SideProject [link] [comments]

OpenAI-compatible API for running open-weight LLMs and video models.
bingus-bongo · HN
Use GLM-5.3 in Cursor today via tokengo API

API providing token-level citations for LLM output grounded in attention analysis.
apoorvumang · HN
TokenPath – token-level citations for LLM output, read from attention

Access multiple AI models through one API with transparent prepaid pricing.
@MyApiTaco · X
GLM 5.3 Flash is now on 🌮⚡ Limited-time promotion: 66.6% OFF retail • Input: $0.05 (retail $0.15) • Cache Input: $0.01 (retail $0.03) • Output: $0.167 (retail $0.50) Plus, get an extra 5% bonus on topups over $100. #GLM #zAI #openrouter #vibecoding

Access 44 AI models from 12 providers through a single unified API.
@Boltchh · X

Track AI coding costs attributed to pull requests, teams, and organizations.
haseebejaz · HN
TokenSpend, the AI ROI Solution

Compare AI model coverage, pricing, uptime, and latency across different AI relays.
zizheruan · HN
XTokenChecker – Verifies model identities of your AI gateway