
Glaux
使用 WebGPU 在浏览器中直接运行 Hugging Face ONNX 模型。
granganath · HN
Glaux – Browser-Only AI for Hugging Face ONNX Community Models
完整作品展
技术栈
60 projects

使用 WebGPU 在浏览器中直接运行 Hugging Face ONNX 模型。
granganath · HN
Glaux – Browser-Only AI for Hugging Face ONNX Community Models

用WebAssembly在浏览器运行和管理LLM模型
userfrom1995 · HN
Goku – WASM (wllama)-powered LLM inference and model manager

压缩提示词以减少向LLM API发送的token数量和成本。
@asgujjuasitgets · X

在 leaderboard 上按官方基准对比 AI 大模型的性能排名
fcten · V2EX
做了一个大模型 leaderboard 网站 最近一个月 CodeX 疯狂送重置,token 根本用不完,顺手做点东西。 地址:[知行录]( https://leaderboard.cn/) 排行依据主要为模型官方基准测试成绩。非主观排名。 数据会持续更新。如果有点用,欢迎各位 v 友收藏~

检测LLM API是否被降智或偷换模型,一键跑6项探针得出结果
cocodot LLM 降智检测 — 免费的 LLM API「降智/偷换模型」在线检测:填入任意 OpenAI 兼容端点的 base_url 和临时 API Key,跑 6 项探针(模型声明、动态题、能力完整性等)生成分项报告;Key 仅用于当次检测、不落库不留存,检测方法[开源](https://github.com/cocodot2026/cocodot-llmprobe)

查看LLM模型在10个基准问题上的评分和排名。
fristovic · HN
She watched me look at model rankings and asked what do the numbers mean... I literally had no good way of explaining it to her so I just came up with something that is approximately in the same ballpark as some of the benchmarks out there lol

压缩LLM提示词和文档以降低token使用和API成本
@marcusyul · X
THEY JUST GAVE AWAY 100 MILLION FREE TOKENS SO YOU CAN STOP BURNING THROUGH YOUR CLAUDE CODE BUDGET. if you code with AI you already know: the session fills up, starts failing, and on top of that you're overpaying there's a tool that fixes this: it shrinks the context before the model even sees it same model, same response, a fraction of the cost in a real session: from $154 to $43. a 72% drop and right now: → extend your Fable sessions in Claude Code → 100M free tokens to try it out you don't switch models you don't touch your code you just stop paying to repeat yourself link below ⬇️

LLM API支出分析仪表板,按模型和环境分类,含优化建议
ATsimbalistov · HN
Show HN: Tracking GenAI cost and endpoint fragility so app teams don't have to

Agent原生的TypeScript框架,在托管GPU上训练和部署定制模型。
@soleil_colza_ · X

用KataGo AI在云GPU上分析围棋游戏
malusama · V2EX
[分享创造] 给没有显卡的棋友:一个可直接接 KaTrain 的云端 KataGo 服务 电脑没显卡,或者笔记本跑不动全盘复盘,但想用 KaTrain 分析棋谱?我做了一个云端 KataGo 服务,KaTrain 里只需要填一个 WSS 地址。 ## 怎么接入 - 注册后拿到 Token: https://go.malu.moe - KaTrain v1.18+ 的设置里,Remote Engine 填:`wss://go.malu.moe/v1/katago/<你的 Token>` - 走的是标准 KataGo Analysis Engine JSON over WebSocket ,复盘、深度分析、Sweep 这些功能都兼容 - 不想用 GUI 的话也有 REST 接口:`POST /v1/analyze` ## 免费额度(注册即用,不用绑卡) - 单次分析最多 1,000 visits - 每月 1,000,000 visits ,按 KaTrain 默认 500 visits/手算,大约够 10 局全盘复盘 - 并发 1 ,最多 5 个 Token

用您的数据微调语言模型并管理自定义事实,获得密码学删除证明。
@MBrew26730 · X
Dataset cleaning + fine tuning + continual learning at

通过现有 API 订阅同时运行多个 AI 模型。
@ContinuumCode · X