
frontier.fast — inference, measured
Submit kernel patches and engine optimizations for LLM inference speed, benchmarked on dedicated hardware.
carsenk · HN
Frontier.fast – Help push the frontier of LLM speed forward
The full gallery
Tech stack
60 projects

Submit kernel patches and engine optimizations for LLM inference speed, benchmarked on dedicated hardware.
carsenk · HN
Frontier.fast – Help push the frontier of LLM speed forward

Visualize hardware performance metrics while running LLM inference on your system.
dev_dan_2 · HN
WatchMachineGo – A visualizer to show hardware performing LLM inference

Real-time LLM-powered news aggregator surfacing trending stories with live updates.
tdubey · HN
DWS A LLM Generated, "Drudge Report" style news site

Run LLM inference on consumer GPUs with NVIDIA TensorRT-LLM optimization.
brianhabana123 · HN
TensorRT-LLM running natively on Windows (no WSL)

Semantic caching reduces LLM token costs and latency for AI queries.
u/ornymo_official · Reddit
how to reduce ai costs there are lots of way to reduce costs but there all complex to setup i know this cause i tried one in production so i built ornymo we cache meaning not the exact string allowing us to give same awnsers thus reducing llm costs and latency check it out at ornymo.com free for a limited time and let me know your feedback submitted by /u/ornymo_official to r/buildinpublic [link] [comments]

Compress prompts and reduce LLM token costs by detecting duplicate tool calls.
@DeveloperL92487 · X
I built my first app in 60min And now I got $500 MRR in one month Check here if you are interested It’s a tool to reduce agent token consumption, speed up agent response, and clean up memory cache

Compare latency and throughput performance across LLM API providers.
@QAInsights · X

Fine-tune LLMs with your data, teach and erase custom facts, get cryptographic deletion proofs.
@MBrew26730 · X
Dataset cleaning + fine tuning + continual learning at

Find AI models optimized for your hardware with performance and pricing estimates.
cdnsteve · HN
Tokenstead, find AI models for your hardware

Compare and evaluate AI models across coding, reasoning, agents, and other benchmarks.
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

Interactive LLM chat interface running on Enclave's confidential compute platform.
SteveDeFacto · HN
Hi HN, I built Enclave, self-serve confidential compute on GPUs. Technical documentation is on the site, but I'd rather show than tell. Here are a couple apps hosted live on the platform: LLM Chat bot: https://cc1f4f3f.app.enclave.host AI Image Generation: https://da09d0f2.app.enclave.host If you have any questions, I would be more than happy to discuss.

View LLM model rankings across 10 benchmark questions.
fristovic · HN
She watched me look at model rankings and asked what do the numbers mean... I literally had no good way of explaining it to her so I just came up with something that is approximately in the same ballpark as some of the benchmarks out there lol