
LLM/trail
Group chat messages by topic to quickly navigate to past conversations.
doantam · HN
I built a chat client that uses embeddings to cluster messages by topic
The full gallery
Tech stack
60 projects

Group chat messages by topic to quickly navigate to past conversations.
doantam · HN
I built a chat client that uses embeddings to cluster messages by topic

View LLM model rankings across 10 benchmark questions.
fristovic · HN
She watched me look at model rankings and asked what do the numbers mean... I literally had no good way of explaining it to her so I just came up with something that is approximately in the same ballpark as some of the benchmarks out there lol

Send your question to a panel of LLMs that peer-review each other and return one synthesized answer.
u/Puzzleheaded-Log-27 · Reddit
Building a multi-model AI deliberation tool taught me something about trust LLM Counsel isn't another wrapper around one model - it sends your question to a panel of frontier LLMs, has them peer-review each other anonymously, and an impartial "chairman" model returns one synthesized answer. Free to start, pay-as-you-go after, credits don't expire. What I've learned so far: people trust a synthesized answer a lot more once they can see that the models actually disagreed and how that disagree

View LLM API spending analytics by model and environment with optimization suggestions.
ATsimbalistov · HN
Show HN: Tracking GenAI cost and endpoint fragility so app teams don't have to

Manage and run LLM models in your browser via WebAssembly.
userfrom1995 · HN
Goku – WASM (wllama)-powered LLM inference and model manager

Route LLM API traffic through a gateway with built-in cost tracking, latency analytics, and PII redaction.
charltonraven · HN
RavenGate – LLM gateway that redacts PII across SSE chunk boundaries

Semantic caching reduces LLM token costs and latency for AI queries.
u/ornymo_official · Reddit
how to reduce ai costs there are lots of way to reduce costs but there all complex to setup i know this cause i tried one in production so i built ornymo we cache meaning not the exact string allowing us to give same awnsers thus reducing llm costs and latency check it out at ornymo.com free for a limited time and let me know your feedback submitted by /u/ornymo_official to r/buildinpublic [link] [comments]

Use one API to access and switch between LLM providers while optimizing inference costs.
justin2025 · Product Hunt
Auriko Trading desk for LLM calls

Transcribe audio and video files to text with AI-generated summaries.
@Mahima_Akkina · X
Yes, you can check - It's a tool used to convert both and video into text and summaries

Anonymous LLM proxy accepting Bitcoin and Monero for API access to Anthropic and OpenAI without an account.
not_wowinter13 · HN
Anonymous LLM proxy. Pay in crypto, no account needed

Monitor AI model calls, agent steps, and retrieval with token tracking, cost analysis, and latency metrics.
ephraimduncan · HN
Observability for Coding Agents and LLM Applications

See the concepts a language model holds at each layer before it answers.
ada1981 · HN
I built a web tool to see and edit what an AI thinks before it answers