
LLM 推理计算器 | LLM Inference Calculator
Estimate GPU memory, latency, TTFT, TPOT, and throughput for LLM inference.
popopanda · HN
LLM Inference Calculator – Estimate VRAM, Latency, and Throughput
The full gallery
Tech stack
60 projects

Estimate GPU memory, latency, TTFT, TPOT, and throughput for LLM inference.
popopanda · HN
LLM Inference Calculator – Estimate VRAM, Latency, and Throughput

Automatically route each prompt to the cheapest capable model to cut API costs.
u/ASDKING100 · Reddit
Launched an AI API router tonight, and found a bug hours in that would've taken real payments without ever upgrading the account Built LLMLite over the past few weeks — it classifies each prompt and routes it to the cheapest model that can actually handle it, instead of hitting GPT-4o for everything. Free tier, no card needed to try it. Tonight, right as I was about to launch, ran a real transaction to test the payment flow end to end. Paddle processed it, webhook fired, signature verified

Compress prompts before LLM API calls to reduce token usage and costs.
@asgujjuasitgets · X

Manage and run LLM models in your browser via WebAssembly.
userfrom1995 · HN
Goku – WASM (wllama)-powered LLM inference and model manager

Host a dedicated LLM instance in the EU with flat-rate pricing and no usage limits.
CodingPanda42 · HN
Virtual Private LLM, fixed fee with no usage or token limits

An LLM gateway for OpenAI, Anthropic, Google and Azure. Every request logged, priced to the token, and audited for waste you can actually recover.
@razdagan3 · X

Visualize hardware performance metrics while running LLM inference on your system.
dev_dan_2 · HN
WatchMachineGo – A visualizer to show hardware performing LLM inference

Convert PDF, DOCX, XLSX, CSV, JSON, XML, HTML, Images to clean Markdown optimized for AI Agents, RAG, and Vector DBs. 100% Privacy-First, In-Browser Conversion.
@13SahajChawla · X
For professionals to redact their client's sensitive informations before giving AI to process it & while converting any kind of document to a structured MD file. Better quality outputs, 100% privacy with on-browser local processing, and fully free!

API providing token-level citations for LLM output grounded in attention analysis.
apoorvumang · HN
TokenPath – token-level citations for LLM output, read from attention

Compress prompts and reduce LLM token costs by detecting duplicate tool calls.
@DeveloperL92487 · X
I built my first app in 60min And now I got $500 MRR in one month Check here if you are interested It’s a tool to reduce agent token consumption, speed up agent response, and clean up memory cache

Grades AI agents' real conversations with an LLM judge, providing A–F scorecards and FinOps analysis.
@tech_maju · X

OpenAI-compatible API for running open-weight LLMs and video models.
bingus-bongo · HN
Use GLM-5.3 in Cursor today via tokengo API