Skip to main content

Few Token Do Trick: The Caveman Skill Topping GitHub Trending

The top trending repo on GitHub today is a Claude Code skill that makes your AI talk like a caveman. It’s a joke. It’s also a real answer to a real bill.

I went looking at what was trending on GitHub this morning and the number one project — nearly 3,000 stars in a single day — is called caveman. Its tagline: “why use many token when few token do trick.” It’s a plugin for Claude Code, Codex, Gemini, Cursor and thirty-odd other agents, and its entire premise is to make your coding assistant drop the filler and answer in terse, grammar-free caveman-speak. Same fix, a third of the words.

It reads like a shitpost. But I just wrote a post about the $12.8B AI-coding economy and how half of GitHub’s commits are now AI-touched — and caveman is a surprisingly sharp comment on that same trend. So let me take the joke seriously for a minute, because there’s a real engineering idea under the grunting.

The thing it’s actually attacking

Every reply your AI agent sends is billed by the token. And the default personality of these models is verbose — “Sure! I’d be happy to help. The issue you’re experiencing is most likely caused by…” Three sentences of throat-clearing before the one line you needed. You pay for every word of that preamble, on every reply, forever.

caveman’s before/after example is the whole pitch:

  • Normal (69 tokens): “The reason your React component is re-rendering is likely because you’re creating a new object reference on each render cycle… I’d recommend using useMemo to memoize the object.”
  • Caveman (19 tokens): “New object ref each render. Inline object prop = new ref = re-render. Wrap in useMemo.”

Identical technical content. A quarter of the tokens. Their benchmarks claim an average 65% output reduction across ten prompts, measured against the Claude API’s own token counts — and crucially, they keep code, commands, and error strings byte-for-byte exact. The compression is applied to the prose, never to the payload.

Why this is cleverer than it looks

The slogan they lead with is the real insight: “Caveman no make brain smaller. Caveman make mouth smaller.” It shrinks what the model says, not what it knows. The reasoning still happens at full fidelity inside the model; only the final rendering to text is compressed. That’s the right place to cut — you’re not asking it to think less, just to stop narrating.

A few details that show it’s more than a one-liner:

  • Levels. lite, full, ultra, and a wenyan mode that renders in classical Chinese — which, only half-jokingly, packs the most meaning per token of any human language.
  • It keeps your language. Write Portuguese, it grunts back in compressed Portuguese. It compresses style, not meaning.
  • Compress the memory file too. /caveman-compress CLAUDE.md rewrites your project’s instruction file into terse form, cutting input tokens on every session after — while preserving code, URLs, and paths verbatim.
  • It measures itself. /caveman-stats reports real session token usage and lifetime dollar savings. The benchmarks live in the repo, committed and reproducible. For a meme project, that’s unusually honest.

Where I’d be careful

I like it, but a token-counter’s enthusiasm shouldn’t switch off the engineer’s skepticism:

  • Output tokens are the cheap half. The savings are on output only — their own chart shows 0% saved on input. In agentic coding, the input side (your files, tool results, the whole context window fed back each turn) is often the bigger cost. caveman helps, but it’s trimming the smaller line item.
  • Terse isn’t always better for humans. When I’m debugging something subtle, the model’s “why” paragraph is sometimes the part that catches its own mistake. Compress the explanation away and you may also compress away the reasoning you’d have caught an error in. For grinding through boilerplate, grunt away. For a tricky design call, I want the full sentence.
  • Curl-pipe-bash installs. The one-line installer pipes a remote script straight into your shell across 30+ agents. Convenient, but it’s exactly the kind of thing I’d read before running — same instinct as any privileged installer.

The takeaway

caveman is a joke with a real spreadsheet behind it. It won’t change how these tools reason, and it won’t touch your biggest cost line — but it’s a genuinely smart observation that the default verbosity of AI agents is a tax you can opt out of, and that you can do it without losing a byte of the technical answer. In a year where we’re all quietly watching our API bills climb, “same brain, smaller mouth” is a better engineering principle than it has any right to be. Few token do trick.

Comments

Popular posts from this blog

Cursor AI Review: Is the AI Code Editor Worth It?

I've been using Cursor as my main code editor for a while now, and enough people have asked whether it's worth switching to that a proper review felt overdue. Short version: for me, yes — but with caveats. What is Cursor? Cursor is an AI-first code editor built as a fork of VS Code. That means every extension, theme, and keybinding you already use in VS Code works here, but with AI woven directly into the editing experience instead of bolted on as a plugin. It's made by Anysphere and can run models from OpenAI and Anthropic under the hood. What I like Tab completion is uncanny. Cursor predicts your next edit — not just the rest of the line, but the next change across the file. Once you get used to hitting Tab, going back to a plain editor feels slow. The Composer / Agent mode. You describe a change in plain language and it edits multiple files at once, showing you a diff to accept or reject. For refactors and boilerplate, this saves real time. It unde...

MacBook Pro M5 vs M5 Pro: Which One Should You Actually Buy?

Apple's latest 14-inch MacBook Pro comes in two very different flavors: the base M5 and the step-up M5 Pro . On paper they look similar — same gorgeous Liquid Retina XDR display, same design — but under the hood the gap is bigger than the names suggest. Here's a clear, no-hype breakdown, with concrete use cases so you can match the chip to your work. Quick spec comparison Spec M5 M5 Pro CPU 10-core (4 performance + 6 efficiency) Up to 18-core (6 performance + 12 efficiency) GPU 10-core Up to 20-core Neural Engine 16-core 16-core Memory bandwidth 153 GB/s 307 GB/s (roughly double) Unified memory 16 / 24 / 32 GB 24 / 48 / 64 GB Max storage Up to 4 TB SSD Up to 8 TB SSD Battery (video playback) Up to 24 hours Up to 22 hours Media engines Single encode/ProRes engine More encode/ProRes engines (higher configs) What actually changes between them More cores — the M5 Pro nearly doubles CPU cores and adds GPU cores, so sustained, multi-threaded work finishe...

Running a Server on a Mac Mini: Apple Silicon vs the Home-Server Field

The Mac Mini has quietly become one of the most interesting home-server boxes you can buy. It’s tiny, nearly silent, sips power, and Apple Silicon punches far above its weight. But is it actually the right machine to run your services on — or are you paying an Apple tax for a job a $400 mini PC does better? Let’s put it head-to-head. Why a Mac Mini makes a surprisingly good server Three things make Apple Silicon compelling as an always-on machine: Performance per watt. This is the headline. An M4 Mini idles at just a few watts and rarely pushes past ~35W under load, while delivering multicore performance that embarrasses machines drawing twice the power. Silence. Under typical server loads the fan is inaudible. If your “server” lives in a living room or bedroom, this matters more than any benchmark. Footprint. It’s the size of a coaster and runs cool, so it tucks anywhere. The honest catch It’s not all upside: macOS isn’t...