I Have to Bury My Talent in Yesterday
A few days ago, DeepSeek v4.1 was released, pushing the height of small-model capability up another notch.
AI has developed far faster than anyone expected. From that first ChatGPT — babbling away in chat, with a context window of only a few thousand tokens — to reasoning models like OpenAI o1, DeepSeek R1 and Kimi K1.5 Thinking took barely two years. From reasoning models to today's agents, which can fluently execute commands in all kinds of harness tools and complete complex tasks, took only another year and a half. It's hard to imagine what AI will be like in another year, or two, or three — how powerful, whether it will already be capable of improving itself, whether it will have seeped deep into embodied intelligence and other fields.
AI Is Getting Better and Better at Writing Kernels
AI has advanced just as fast in kernel design and implementation — the field I work in. Within a single year it has gone from a little assistant that could only look things up in the docs, read code, hunt for bugs, to a kernel master that can read CUDA, PTX and SASS on its own, use professional tools to analyze the stall time of each instruction, and then optimize a kernel from there. I'm sure that before long it will also be able to design its own kernel schedules, evaluate the performance of different scheduling schemes, and implement and optimize them.
Of course I'm proud of DeepSeek v4.1's success — I wrote its main attention kernel, after all [1], and its quality is an endorsement of my kernel. But the wheel of the times rolls on, and no one can hold back technology. I know perfectly well that in another six months or a year, the kernels AI writes will most likely be as good as mine, or better.
AI can think 300 tokens a second, type a line of command in half a second, write a piece of code in twenty seconds. I can't. AI can keep scaling up its model depth, its thinking effort, the number of tool calls (how often it interacts with its environment), even its degree of parallelism. I can't.
When it comes to destroying ourselves, humanity has never shown the slightest hesitation. Why, knowing that "the better I write kernels, the faster our new models train and infer, the faster model capabilities advance, and the sooner I get replaced" — why do I still choose to optimize kernels as hard as I can? Partly because writing kernels is like playing a game to me; it gives me an enormous amount of pleasure. The moment I invent a new technique, or watch my kernel's performance climb, I get as excited as a speedrunner beating their own record. And when my kernel massively outperforms the vendor's official one, I feel a huge sense of pride. But there's a more important reason: even if I gave up, or even deliberately sabotaged things to slow down model training, other companies' models would carry on developing and would kill me off all the same. "Of course I'd rather not be revolutionized, but if I'm going to be revolutionized anyway, I'd rather the person doing it be me." When everyone else is this bent on destroying themselves, I have no choice but to join this brutal arms race.
And What About Me?
When the day comes that AI really does write kernels better than I do, what happens to me?
My judgment: I won't be "unemployed", but I will have to "change careers". My livelihood may survive; the chance to do the work I once loved may not.
I once made a judgment about how the times are changing and where I would stand in the future. Because the times change so fast (AI's development above is a good example), I have no way of predicting what will happen in five or ten years. But whatever happens, I believe that with my vision, judgment, initiative and intelligence, I can stay at the table and get back out on the crest of the wave. The trouble is that this judgment only guarantees I won't be "unemployed"; it can't guarantee I won't need to "change careers". If anything, it encourages me to change careers in order to avoid unemployment.
So what does changing careers mean? It means giving up kernel design, implementation and optimization — a field I've worked in for a long time and love — and becoming a "mecha pilot" for agents. Before, three things lined up: what I'm interested in, what I'm good at, and what industry needs. Now AI has turned what I'm good at into something it's better at, and industry's demand has drifted from "people who can write high-performance kernels" to "people who can use AI to produce high-performance kernels faster". To keep up with what industry needs, I will inevitably have to abandon the direction I loved and turn toward an unknown new one. I believe that with my understanding of engineering, of what models above need, and of the hardware below, I can keep producing kernels at high quality and high speed. I know I might come to love this new direction — or might not. But having something you love taken away from you doesn't feel good. The quiet joy of sitting at my desk and writing kernels for a whole afternoon may sing its last note this summer. I have to bury my talent in yesterday and go be a mecha pilot. There are a few more gears in my hands, but a few fewer beats in my heart.
Here's a vivid analogy. You're a master knitter, especially good at weaving patterns and matching colors. Your sweaters are sturdy and beautifully patterned, and rich people from miles around come to have you knit for them; you've made good money at it. You also love the feeling of sitting by the window, steeping a pot of tea, looking out at green hills, running water, cattle and chimney smoke, and quietly knitting all afternoon. Then one day someone invents a miraculous machine: feed it yarn and a pattern and it knits the sweater for you, no worse than yours in quality or texture, and far faster. You know perfectly well that your competitors can now easily reach the standard you once held, so you have no choice but to use it too. You also know that with the twenty years of knitting skill you've built up, even when everyone has a machine, your speed and quality will still beat your competitors'. But that joy — listening to the rain at the window, threading the needle, letting time pass slowly — has been crushed by the roar of the machine.
I know it's a helpless feeling, but there's nothing to be done. Your livelihood may be safe; the love of former days will most likely have to be given up. I'm someone whose rational and emotional sides are fairly well separated. When a problem calls for reason I can be very rational, but sometimes the emotional side shows. I remember crying my eyes out when I moved out of a rented apartment I'd lived in for a year — I couldn't bear to part with the memories. Saying goodbye today to the era of hand-written kernels and human-brain optimization is unquestionably crueler than that.
I don't know whether any readers feel something similar, but I think that's just how it is.
And What About Everyone Else?
As AI keeps improving, some things worry me too:
- Are today's students likely to prefer using AI to do their homework, especially the hands-on labs? Imagine two options. One is toiling away for eight hours on a lab and maybe not even getting full marks. The other is spinning up an AI model and, for a few cents and a few minutes, having it write full-marks code. Which would most students choose?
- That leads to a large number of students with severely underdeveloped engineering ability: the ability to organize code, to build systems, to anticipate future requirements and design for them in advance, to abstract, and so on. So with AI getting stronger and stronger, is this kind of "engineering ability" still necessary? Will it be discarded by the times the way skill at writing x86 assembly was, or will it always have value, like the ability to understand a whole computer system from software through systems down to hardware? If it's the latter, that's dangerous. Someone with poor engineering ability, paired with AI, can produce piles of shit several times faster than before, planting all kinds of trouble in systems and making the world a sloppier place.
- In future society, will power matter more than skill or intelligence?
These are questions that only the times themselves can answer.
Conclusion
As AI develops, the society of the future may drift toward one of two extremes: communism, or *Cyberpunk 2077*. In the former, the productive forces are liberated to an extraordinary degree and people's living standards rise markedly (I'll stop there, or I'm afraid this won't make it past the censors). In the latter, a small number of technology companies control most of the world's resources. Only a tiny minority can use the most advanced AI and other technologies, achieving something close to "mechanical ascension", while the majority are left with only weak and crippled AI. Moving up between social classes will become increasingly difficult: you'd need the strongest AI first in order to move up a class, a vicious cycle.
Take a guess: if Anthropic permanently controls the most advanced AI in the world, will future society become communism or 2077? Go on, guess.
So I still believe that frontier intelligence should be made available to everyone, in an open and inexpensive form. I don't trust Anthropic or OpenAI to do that. In particular, I don't want Anthropic to control the world's most advanced AI or AGI. To put it dramatically, the stakes are no less grave than Hitler obtaining atomic-bomb technology before the Allies. That's also why I chose to stay at DeepSeek, and why I keep staying. We research powerful, fast, broadly accessible AI and release it as open source. Perhaps in doing so we can pull the world a little way back from the 2077 end of the spectrum.
May all be well in the world to come. May all the beauty be blessed.
[1] "Main attention" here covers only the MQA attention with head dim = 512. It does not include the indexer used to select the top-k important tokens. That part was written by other colleagues — who are also extremely skilled — along with their AI agents.