
BareMetalRT — Bare Metal AI
用NVIDIA TensorRT-LLM在消费级GPU上进行高性能大语言模型推理。
brianhabana123 · HN
TensorRT-LLM running natively on Windows (no WSL)
完整作品展
技术栈
60 projects

用NVIDIA TensorRT-LLM在消费级GPU上进行高性能大语言模型推理。
brianhabana123 · HN
TensorRT-LLM running natively on Windows (no WSL)

查看 GPU 兼容的 LLM 模型及成本预算方案。
jaeseok614 · HN
Open-source calculator for "will my GPU run this LLM?"

实时可视化硬件在运行LLM推理时的性能指标
dev_dan_2 · HN
WatchMachineGo – A visualizer to show hardware performing LLM inference

估算llama.cpp模型推理所需的显存大小
hypfer · HN
According to this shitty vibecoded thing "I" built https://hypfer.github.io/will-it-fit-llama-cpp/ (and I guess according to math too), FP16 K/V would give me something like 90k context at the same model quant, which doesn't really fit my usage. But maybe someone else has experience to share there

在 30 秒内测量 GPU 的真实内存带宽上限,用于本地 AI。
Ar5en1c · HN
Headroom – measure your GPU's true bandwidth ceiling for local AI

查找与您硬件兼容的AI模型并查看性能和价格估计。
cdnsteve · HN
Tokenstead, find AI models for your hardware

提交 LLM 推理优化内核,在专用硬件上进行基准测试并竞争排名。
carsenk · HN
Frontier.fast – Help push the frontier of LLM speed forward

Runpod是无服务器GPU推理平台,冷启动低于200ms,按秒计费。
@svpino · X
You can check out Runpod here: Thanks to the Runpod team for partnering with me on this post.

在欧盟托管私有 LLM 实例,固定月费无使用限制。
CodingPanda42 · HN
Virtual Private LLM, fixed fee with no usage or token limits

The compatibility engine for local AI. Tell us what you want to run — we'll tell you exactly which models fit your machine, with verified benchmarks.
@Carl0sFelipe · X
Just shipped — a tool that helps you discover which local AI models actually run on your hardware, with community benchmarks, quantization support, and estimated speed. Building in public from here. #BuildingPublic #AIDevelopment #rust #benchmaks #aimodel

在浏览器中运行AI模型基准测试以检测性能回归。
pepperpoppins · HN
Trunchbull, run real models against any benchmark in your browser

追踪您的应用中使用的 AI 模型,并在其被弃用前获得警告。
taylorgt · HN
Find every AI model your code calls and warn before it's retired