
FlexInference: Drop your AI costs today
Route your LLM API requests across multiple providers to cut costs and meet latency targets.
Aperswal · HN
Made a Free LLM Router
The full gallery
Tech stack
21 projects

Route your LLM API requests across multiple providers to cut costs and meet latency targets.
Aperswal · HN
Made a Free LLM Router

Route LLM calls to cost-effective models without sacrificing quality.
george_avila · Product Hunt
IQ Routing Trajectory-aware LLM routing that cuts agent cost

Automatically route each prompt to the cheapest capable model to cut API costs.
u/ASDKING100 · Reddit
Launched an AI API router tonight, and found a bug hours in that would've taken real payments without ever upgrading the account Built LLMLite over the past few weeks — it classifies each prompt and routes it to the cheapest model that can actually handle it, instead of hitting GPT-4o for everything. Free tier, no card needed to try it. Tonight, right as I was about to launch, ran a real transaction to test the payment flow end to end. Paddle processed it, webhook fired, signature verified

Format-agnostic LLM hub. Bring your own provider keys and route across Anthropic, OpenAI, ChatGPT/Codex, Kimi, Alibaba DashScope, and AWS Bedrock — with unified observability and c
@0xxmemo · X

Route requests to Claude, GPT, Codex and more through a single OpenAI-compatible API.
@RouteraOne · X
daily limits turning vibe coding into handless mode 😭 Routera gives you usage-based access to Claude, Codex, GPT and more, without daily or weekly caps, and it’s usually cheaper than stacking subscriptions

An LLM gateway for OpenAI, Anthropic, Google and Azure. Every request logged, priced to the token, and audited for waste you can actually recover.
@razdagan3 · X

Use one API to access and switch between LLM providers while optimizing inference costs.
justin2025 · Product Hunt
Auriko Trading desk for LLM calls

Tests LLM endpoints with adversarial cases and provides OWASP-mapped security audit reports.
@aryaan_sheth · X
- LLM security for small teams

Semantic caching reduces LLM token costs and latency for AI queries.
u/ornymo_official · Reddit
how to reduce ai costs there are lots of way to reduce costs but there all complex to setup i know this cause i tried one in production so i built ornymo we cache meaning not the exact string allowing us to give same awnsers thus reducing llm costs and latency check it out at ornymo.com free for a limited time and let me know your feedback submitted by /u/ornymo_official to r/buildinpublic [link] [comments]

Routes browser tasks across 250+ runner combinations with one OpenAI-compatible API.
ygabriel27 · HN
Banana Peel – OpenRouter for browser agents

OpenAI-compatible API for running open-weight LLMs and video models.
bingus-bongo · HN
Use GLM-5.3 in Cursor today via tokengo API

Detect if your LLM API has been model-swapped or degraded with 6 deterministic probes.
cocodot LLM 降智检测 — 免费的 LLM API「降智/偷换模型」在线检测:填入任意 OpenAI 兼容端点的 base_url 和临时 API Key,跑 6 项探针(模型声明、动态题、能力完整性等)生成分项报告;Key 仅用于当次检测、不落库不留存,检测方法[开源](https://github.com/cocodot2026/cocodot-llmprobe)