
BareMetalRT — Bare Metal AI
Run LLM inference on consumer GPUs with NVIDIA TensorRT-LLM optimization.
brianhabana123 · HN
TensorRT-LLM running natively on Windows (no WSL)
The full gallery
Tech stack
60 projects

Run LLM inference on consumer GPUs with NVIDIA TensorRT-LLM optimization.
brianhabana123 · HN
TensorRT-LLM running natively on Windows (no WSL)

Check which LLMs run on your GPU given your hardware specs and budget.
jaeseok614 · HN
Open-source calculator for "will my GPU run this LLM?"

Visualize hardware performance metrics while running LLM inference on your system.
dev_dan_2 · HN
WatchMachineGo – A visualizer to show hardware performing LLM inference

Estimate VRAM requirements for running models with llama.cpp
hypfer · HN
According to this shitty vibecoded thing "I" built https://hypfer.github.io/will-it-fit-llama-cpp/ (and I guess according to math too), FP16 K/V would give me something like 90k context at the same model quant, which doesn't really fit my usage. But maybe someone else has experience to share there

Measure your GPU's real memory-bandwidth ceiling for local AI in 30 seconds.
Ar5en1c · HN
Headroom – measure your GPU's true bandwidth ceiling for local AI

Find AI models optimized for your hardware with performance and pricing estimates.
cdnsteve · HN
Tokenstead, find AI models for your hardware

Submit kernel patches and engine optimizations for LLM inference speed, benchmarked on dedicated hardware.
carsenk · HN
Frontier.fast – Help push the frontier of LLM speed forward

Deploy AI inference models on serverless GPUs with sub-200ms cold starts and pay-per-second billing.
@svpino · X
You can check out Runpod here: Thanks to the Runpod team for partnering with me on this post.

Host a dedicated LLM instance in the EU with flat-rate pricing and no usage limits.
CodingPanda42 · HN
Virtual Private LLM, fixed fee with no usage or token limits

The compatibility engine for local AI. Tell us what you want to run — we'll tell you exactly which models fit your machine, with verified benchmarks.
@Carl0sFelipe · X
Just shipped — a tool that helps you discover which local AI models actually run on your hardware, with community benchmarks, quantization support, and estimated speed. Building in public from here. #BuildingPublic #AIDevelopment #rust #benchmaks #aimodel

Run real models against benchmarks in your browser to detect performance regressions before production.
pepperpoppins · HN
Trunchbull, run real models against any benchmark in your browser

Track AI models used in your apps and receive warnings before they're deprecated.
taylorgt · HN
Find every AI model your code calls and warn before it's retired