
LLM 推理计算器 | LLM Inference Calculator
Estimate GPU memory, latency, TTFT, TPOT, and throughput for LLM inference.
popopanda · HN
LLM Inference Calculator – Estimate VRAM, Latency, and Throughput
The full gallery
Tech stack
23 projects

Estimate GPU memory, latency, TTFT, TPOT, and throughput for LLM inference.
popopanda · HN
LLM Inference Calculator – Estimate VRAM, Latency, and Throughput

Compare LLM API pricing across 450+ routes and calculate real monthly costs with caching and batch pricing factors.
u/Greywolff06 · Reddit
I built LLMPrice — a free calculator for comparing LLM API costs across 450+ pricing routes I kept running into the same problem when comparing LLM APIs: the headline token price doesn't always tell you what your actual workload will cost. Caching, batch pricing, reasoning tokens, retries, different endpoints, and OpenRouter routes can change the result quite a bit. So I built LLMPrice.com. You enter your workload once — requests, input/output tokens, caching, retries, etc. — and it com

Convert PDF, DOCX, XLSX, CSV, JSON, XML, HTML, Images to clean Markdown optimized for AI Agents, RAG, and Vector DBs. 100% Privacy-First, In-Browser Conversion.
@13SahajChawla · X
For professionals to redact their client's sensitive informations before giving AI to process it & while converting any kind of document to a structured MD file. Better quality outputs, 100% privacy with on-browser local processing, and fully free!

API providing token-level citations for LLM output grounded in attention analysis.
apoorvumang · HN
TokenPath – token-level citations for LLM output, read from attention

Grades AI agents' real conversations with an LLM judge, providing A–F scorecards and FinOps analysis.
@tech_maju · X

Compare latency and throughput performance across LLM API providers.
@QAInsights · X

Semantic caching reduces LLM token costs and latency for AI queries.
u/ornymo_official · Reddit
how to reduce ai costs there are lots of way to reduce costs but there all complex to setup i know this cause i tried one in production so i built ornymo we cache meaning not the exact string allowing us to give same awnsers thus reducing llm costs and latency check it out at ornymo.com free for a limited time and let me know your feedback submitted by /u/ornymo_official to r/buildinpublic [link] [comments]

@launch_llama https://t.co/YsRI7cTzqm https://t.co/EXOr5fEkp0
@melonrice383235 · X


Use one API to access and switch between LLM providers while optimizing inference costs.
justin2025 · Product Hunt
Auriko Trading desk for LLM calls

@launch_llama An movie platform https://t.co/WnbiDZbHH8 https://t.co/LWyGNRUUFD
@Kelvinwz5lma · X
An movie platform