
InferAll — Unified AI API for OpenAI, Claude, Gemini & 207+ Models
Access 207+ AI models from different providers through a single unified API.
TaylorM492 · HN
InferAll – One API for OpenAI, Anthropic, Google, Nvidia Nim
The full gallery
Tech stack
60 projects

Access 207+ AI models from different providers through a single unified API.
TaylorM492 · HN
InferAll – One API for OpenAI, Anthropic, Google, Nvidia Nim

Detect if your LLM API has been model-swapped or degraded with 6 deterministic probes.
cocodot LLM 降智检测 — 免费的 LLM API「降智/偷换模型」在线检测:填入任意 OpenAI 兼容端点的 base_url 和临时 API Key,跑 6 项探针(模型声明、动态题、能力完整性等)生成分项报告;Key 仅用于当次检测、不落库不留存,检测方法[开源](https://github.com/cocodot2026/cocodot-llmprobe)

Compress prompts and reduce LLM token costs by detecting duplicate tool calls.
@DeveloperL92487 · X
I built my first app in 60min And now I got $500 MRR in one month Check here if you are interested It’s a tool to reduce agent token consumption, speed up agent response, and clean up memory cache

Track AI models used in your apps and receive warnings before they're deprecated.
taylorgt · HN
Find every AI model your code calls and warn before it's retired

Access thousands of AI models through a single OpenAI-compatible API.
@mageofweb3 · X

Use a single API to access AI models from 12 different providers.
@Boltchh · X

Manage API keys and usage budgets for coding-agent workflows with request routing.
u/Zyron_X · Reddit
I built a service for people to use Codex API without 5-hour limit disruption I built a small service for people who use the OpenAI Codex API regularly and want more predictable usage without the 5-hour or weekly limits. It currently provides: Frontier OpenAI models (GPT 5.6 family included) Managed API key Monthly usage budgets depending to plan No 5-hour limit No weekly limit Under the hood, it is built on top of an open-source project and proxies requests to

Generate MCP endpoints and llms.txt files from any website for AI agents.
AshHackerNews · HN
AgentReady – MCP server that makes any docs site queryable by AI agents

Unified API gateway providing OpenAI, Claude, and Gemini-compatible access to multiple AI models.
@YinsenW_ · X
CherryIN 平台已经上线了DeepSeek v4 flash Vision-Exp 模型,为全球用户带来最领先的多模态AI 服务哈哈哈

Semantic caching reduces LLM token costs and latency for AI queries.
u/ornymo_official · Reddit
how to reduce ai costs there are lots of way to reduce costs but there all complex to setup i know this cause i tried one in production so i built ornymo we cache meaning not the exact string allowing us to give same awnsers thus reducing llm costs and latency check it out at ornymo.com free for a limited time and let me know your feedback submitted by /u/ornymo_official to r/buildinpublic [link] [comments]

Compare LLM API pricing across 450+ routes and calculate real monthly costs with caching and batch pricing factors.
u/Greywolff06 · Reddit
I built LLMPrice — a free calculator for comparing LLM API costs across 450+ pricing routes I kept running into the same problem when comparing LLM APIs: the headline token price doesn't always tell you what your actual workload will cost. Caching, batch pricing, reasoning tokens, retries, different endpoints, and OpenRouter routes can change the result quite a bit. So I built LLMPrice.com. You enter your workload once — requests, input/output tokens, caching, retries, etc. — and it com

Access multiple AI models through one API with transparent prepaid pricing.
@MyApiTaco · X
GLM 5.3 Flash is now on 🌮⚡ Limited-time promotion: 66.6% OFF retail • Input: $0.05 (retail $0.15) • Cache Input: $0.01 (retail $0.03) • Output: $0.167 (retail $0.50) Plus, get an extra 5% bonus on topups over $100. #GLM #zAI #openrouter #vibecoding