
BareMetalRT — Bare Metal AI
Run LLM inference on consumer GPUs with NVIDIA TensorRT-LLM optimization.
brianhabana123 · HN
TensorRT-LLM running natively on Windows (no WSL)
The full gallery
Tech stack
60 projects

Run LLM inference on consumer GPUs with NVIDIA TensorRT-LLM optimization.
brianhabana123 · HN
TensorRT-LLM running natively on Windows (no WSL)

Single API key to access Claude, GPT, and GLM models with $25 daily free credits.
@Awais_209 · X
Claude Opus 4.8, GPT-5.5 & GLM-5.2 for free. Get $25/day in credits. No trial or waitlist. One API key works with Claude Code, Cline, Cursor, Roo & OpenAI-compatible tools. Try it: #FreeTier #ClaudeAPI #CodingTools @AgentRouter_0

Send your question to a panel of LLMs that peer-review each other and return one synthesized answer.
u/Puzzleheaded-Log-27 · Reddit
Building a multi-model AI deliberation tool taught me something about trust LLM Counsel isn't another wrapper around one model - it sends your question to a panel of frontier LLMs, has them peer-review each other anonymously, and an impartial "chairman" model returns one synthesized answer. Free to start, pay-as-you-go after, credits don't expire. What I've learned so far: people trust a synthesized answer a lot more once they can see that the models actually disagreed and how that disagree

Compress prompts before LLM API calls to reduce token usage and costs.
@asgujjuasitgets · X

Route API requests to OpenAI, Anthropic, Gemini, and other AI models using a single API key.
@AstroBo71280348 · X
Hi , I build for indian developers to access all ai llm models on one platform,it also supports UPI payment no visa credit card required,It is more transferent and better then openrouter

Generate and compare offers, emails, and ad copy across seven AI models.
@MarkZofMarkZ · X
Let's connect! Profit Router 📈

Manage API keys and usage budgets for coding-agent workflows with request routing.
u/Zyron_X · Reddit
I built a service for people to use Codex API without 5-hour limit disruption I built a small service for people who use the OpenAI Codex API regularly and want more predictable usage without the 5-hour or weekly limits. It currently provides: Frontier OpenAI models (GPT 5.6 family included) Managed API key Monthly usage budgets depending to plan No 5-hour limit No weekly limit Under the hood, it is built on top of an open-source project and proxies requests to

Monitor AI model calls, agent steps, and retrieval with token tracking, cost analysis, and latency metrics.
ephraimduncan · HN
Observability for Coding Agents and LLM Applications

Compress LLM prompts and documents to reduce token usage and API costs.
@marcusyul · X
THEY JUST GAVE AWAY 100 MILLION FREE TOKENS SO YOU CAN STOP BURNING THROUGH YOUR CLAUDE CODE BUDGET. if you code with AI you already know: the session fills up, starts failing, and on top of that you're overpaying there's a tool that fixes this: it shrinks the context before the model even sees it same model, same response, a fraction of the cost in a real session: from $154 to $43. a 72% drop and right now: → extend your Fable sessions in Claude Code → 100M free tokens to try it out you don't switch models you don't touch your code you just stop paying to repeat yourself link below ⬇️

Add form handling to any website with validation, spam blocking, and intelligent lead routing—no backend required.
@evcodebr · X

Visualize and simulate the BBRv3 congestion control algorithm in your browser.
dilyevsky · HN
BBRv3 for gVisor's netstack, visualized in the browser using WASM

API providing token-level citations for LLM output grounded in attention analysis.
apoorvumang · HN
TokenPath – token-level citations for LLM output, read from attention