Everyone expected the deepseek harness vs claude code test to end with me switching.
It did not.
DeepSeek Harness beat Claude Code on speed by nearly three to one, and on cost by about 57 to one.
I still opened Claude Code the next morning, and so did Kasra Dash, who has used it every single day since it launched.
That sounds irrational until you look at what came out of each one.
This post is the honest version of why the cheaper, faster tool did not win my workflow.
The short answer
Claude Code knows things about my business that I never put in the prompt.
DeepSeek Harness starts from zero every time.
On a one-off script that difference does not matter.
On the twentieth page of a client site, it is the whole job.
What we tested, exactly
One prompt, sent to both agents, with no extra context and no follow-up messages.
The ask was a 3D animated accountancy website, plus a simple game.
DeepSeek Harness ran DeepSeek V4 Pro.
Claude Code ran Claude Opus 5 on high.
Frontier model against frontier model, so nobody can say we handicapped either side.
Where DeepSeek genuinely won
It finished in 11 minutes, start to actual finish.
Claude was still building at 30 minutes and had not stopped when we moved on.
The DeepSeek run cost about 5 cents out of a $10 top-up.
And the harness itself is free, open source, and hit 105,000 GitHub stars in roughly two days — one of the fastest growing open source projects anybody has tracked.
Those are not small wins.
If somebody handed me that result a year ago I would have called it impossible.
Where Claude quietly won
Open the two homepages next to each other and you see it before you read a word.
The Claude build has animation that follows your mouse, so the page feels alive rather than decorated.
It also mentioned Northwest England.
Nobody typed that into the prompt.
Claude had context from earlier work and used it, which is exactly what you want from something that is meant to act like a team member.
The DeepSeek version came out animated and cartoony.
Its Tetris worked, and it was fine.
But a cartoon game on an accountancy website is a mismatch, and any client would spot it in a second.
The Claude game simply felt smoother.
The scoreboard
| What matters | DeepSeek Harness | Claude Code |
|---|---|---|
| Finished the build in | 11 minutes | Over 30 minutes |
| Used unprompted business context | No | Yes |
| Output felt client-ready | No | Closer, still not ready |
| Tokens for the same job | 483,000 | 48,000 at 20 minutes |
| Relative cost per token | ~57x cheaper | Baseline |
| Our score | 7 / 10 | 9 / 10 |
| Maturity | v0.1 preview | Mature |
Neither of them was actually finished
This is the part people skip when they post the screenshots.
Neither build was ready to go live.
Both needed more back and forth, more context, and better prompting before they were worth showing anybody.
So the real comparison is not "which one is done".
It is "which one gets me to done with fewer rounds".
Today that is Claude, because it starts closer and it remembers more.
The token problem I keep pointing at
DeepSeek burned 483,000 tokens on a single one-page website.
Claude was on 48,000 at the twenty-minute mark for the same prompt.
That is a very verbose model, and verbosity is only free while the price is low.
At roughly a 57th of Claude's cost, DeepSeek still wins the maths comfortably.
But run it long enough, or wait for a price change, and that gap narrows in a way most people have not modelled.
I would rather build my workflow on output quality than on somebody else's pricing decision.
What I am actually moving across
I am not being precious about this.
Volume work belongs on the cheap engine.
Drafts, scaffolding, throwaway experiments, anything where I expect to bin the first three attempts — DeepSeek Harness is perfect for that, because attempts cost pennies.
Client-facing builds stay on Claude.
And I do not switch by hand.
I let the agent operating system act as the orchestrator and delegate the job to whichever engine fits.
I did not even install DeepSeek Harness myself — Claude set it up, tested it, and wired it in, which meant I never had to learn another interface.
🔥 Want the setup that routes work to the cheapest engine automatically?
That is exactly what I build inside the AI Profit Boardroom — 4,000+ members, weekly coaching calls, and the full Agent OS walkthrough.
The case for DeepSeek getting there
I gave the harness a 7 out of 10 and Kasra thought I was being generous.
Here is my reasoning.
It is version 0.1. A developer preview. Roughly a tenth of what it will be at a real release.
It is free, it is open, and the community is already all over it.
And competition is the thing that actually moves these products.
If DeepSeek ships something revolutionary in two or three months, Claude has to respond.
That is worth more to you than whichever tool is marginally ahead this week.
FAQ
Did DeepSeek Harness beat Claude Code? On speed and cost, comfortably. On the quality of the finished build, no — the Claude version looked more professional and used context it was never given.
How long did each one take? DeepSeek Harness finished in 11 minutes. Claude Code was still running past 30 minutes on the identical prompt.
Which models were used? DeepSeek V4 Pro in the harness, Claude Opus 5 on high in Claude Code. Like for like.
Is DeepSeek Harness worth installing? Yes, if you have volume work. It is free, it is open source, and at around 5 cents a build the cost of trying it is nothing.
What is the smartest way to run both? Put an orchestrator in front of them. Let it send cheap, high-volume jobs to DeepSeek Harness and client-facing work to Claude, so you never have to choose manually.
What 105,000 stars in two days actually tells you
DeepSeek Harness hit 105,000 GitHub stars in roughly two days.
That is one of the fastest growing open source projects anyone has tracked, and it is worth being precise about why.
It is not because the output beat Claude. It did not.
It is because the harness is free, open, and completely model-agnostic.
You run the agent locally, and you choose the brain that goes inside it.
If you want to pay nothing at all, you can plug a free model in — OpenCode works as a free brain inside the harness — and your running cost drops to zero.
That is a different category of product to a closed tool with a monthly fee and one fixed model.
The stars are not a quality score.
They are a vote on ownership.
And that vote is the thing Claude actually has to answer, more than any single benchmark.
The v0.1 asterisk on every criticism
Everything negative in this post comes with the same footnote.
This is a developer preview. Version 0.1.
Roughly a tenth of what it is likely to be at a real release.
Judging it as a finished product is like reviewing a house at the foundation stage and complaining about the paint.
That is why I landed on a 7 out of 10 when Kasra scored it lower.
He was rating what is in front of him today, which is fair.
I was rating the trajectory, which is also fair.
Both of those can be true, and if you are deciding whether to spend an hour on it this week, the trajectory matters more than the current polish.
The part I actually care about: competition
Here is the honest reason I am glad this launched.
For most of the last year, if you wanted a serious coding agent you had one obvious answer and you paid whatever it cost.
Now there is a free, open, fast alternative with a frontier model behind it and a community shipping plugins for it.
That does not have to beat Claude to matter.
It only has to be good enough that Claude cannot stand still.
If DeepSeek ships something genuinely revolutionary in two or three months, everything gets better, including the tool I am defending in this post.
That is the whole argument for taking an afternoon to try it even if you are staying where you are.
You are not evaluating a replacement.
You are building the option to leave, which is the thing that keeps your costs honest.
One more thing the test exposed
Watching both agents run back to back showed something neither scoreboard captures.
Claude Code did not just produce a better page — it produced a page that understood the brief.
An accountancy website is a tone before it is a layout, and the cartoon game DeepSeek bolted on proved it had matched the words in the prompt without matching the intent behind them.
That is the gap that context closes, and context is the thing you build up over months inside one tool.
It is also the argument for not scattering your work across five agents at random.
Pick where your context lives, keep it there, and treat the other engines as workers rather than homes.
That way you get the cost savings without fragmenting the one asset that actually compounds: everything your primary agent already knows about your business.
Neither build was finished, remember.
Both needed more prompting and more context before they were worth showing anybody.
The agent that starts with more of your context needs fewer of those rounds, and fewer rounds is the real productivity number — not wall-clock speed.
About Julian
I'm Julian Goldie — AI entrepreneur, SEO expert, and founder of the AI Profit Boardroom (4,000+ members). I help business owners scale with AI agents, automation, and SEO.
- 400K+ YouTube subscribers
- 7-figure AI agency (Goldie Agency)
- Daily training inside the Boardroom
- Author of multiple AI automation playbooks
→ Get my best AI training inside the AI Profit Boardroom
Related reading
That is the honest deepseek harness vs claude code answer: the cheap one earned a place in my stack, not the top of it.











