华为廖恒谈超级AI计算硬件(全文转载 + 中文结构化解读)
华为 Fellow 廖恒(半导体首席科学家)4.5 小时深度访谈英文全译转载,附中文结构化解读:十八层宝塔、Tau 定律、Cube/Vector 算力比、昇腾 910→950 与协同设计。
来源标注(全文转载)
- 文本原文(英文全译):Hamish Low, Transcript of an interview with Huawei’s Liao Heng, Cambrian(Substack), 2026-08-04. 原文链接:https://cambrianr.substack.com/p/transcript-of-an-interview-with-huaweis
- 视频来源(中文原声首发):张小珺商业访谈录(语言即世界工作室), Bilibili, 2026-07-25. 视频链接:https://www.bilibili.com/video/BV1nB3u6tERu/
- 说明:上述文本为译者基于 Bilibili 中文访谈的英文转写翻译(claude code + Whisper large-v3),属 AI 转录产物,个别术语/数字或有偏差,请以视频原声为准。本页仅作全文转载与结构化解读,著作权归原作者与首发方所有。
第一部分:正文全文转载(Original Article,英文)
Transcript of an interview with Huawei’s Liao Heng
Below is an interview with Dr. Liao Heng (廖恒), a Huawei Fellow and the company’s Chief Scientist for semiconductors. In July 2026 he sat down with Zhang Xiaojun (张小珺) for a four and a half hour interview. A fairly rare long public address by a leading figure within Huawei’s AI chip efforts.
The interview was published on Bilibili by 张小珺商业访谈录 (Zhang Xiaojun’s Business Interview series, produced by the 语言即世界 studio) on 25 July 2026. All credit for the interview belongs to them. I simply struggled to find an English translation and so produced my own which I’m sharing here.
The method flowed entirely through claude code, with it downloading the video, transcribing with Whisper large-v3 and then cleaning, translating and lightly editing for readability. I had it cross-check against another independent Chinese transcription and press coverage, and found no issues, but keep in mind this is pure AI transcription and translation so some errors are likely.
Prologue
Zhang Xiaojun: Hello everyone, I’m Xiaojun. Today our guest is Dr. Liao Heng (廖恒), Huawei Fellow and Chief Scientist for semiconductors. This should be the first time since Huawei’s ordeal of 2020 that a Huawei executive has come out to talk about how Huawei’s Ascend (昇腾) chips climbed step by step out of the trough. At the same time, this is also my own first time learning about the chip and semiconductor industry. So on one hand we’ll talk about the history of Ascend and China’s choices and story; on the other, from a more macro perspective, we’ll talk about the history, patterns, and overall picture of the entire global semiconductor industry. If you like our video, please give it a like, coin, and favorite. What follows is my interview with Dr. Liao Heng.
What changed in your state of mind? The difference when designing the 910 versus the 950 — the two generations of chips before and after the supply cut-off.
Liao Heng: This change in state of mind — I think putting it into words may sound a bit… What I want to say is: try to imagine, if you were Dong Cunrui, about to hold up that explosive charge — right? — ordinary people like us simply cannot imagine what his state of mind was. Or the soldiers at Shangganling (Triangle Hill) — what was their state of mind? Or — I once saw an interview with a Chinese soldier during the war, which is also a story from Huawei’s publicity posters. A reporter asked him: what do you want once the fighting is over? He said: I don’t think about that, because my parents are dead too… no point thinking about those things. So I’d say the so-called change in state of mind no longer has much meaning. It’s just: you face a difficulty, and you want to solve it.
Zhang Xiaojun: Today Dr. Liao made one request of me: he doesn’t want the focus on him personally. But he has witnessed more than 30, even 40 years of the global chip and semiconductor industry’s development, so I’d like to enter from that angle — through your eyes, take us traveling into this stretch of semiconductor industry history. Setting personal experience aside and talking about the industry itself: over these past 30–40 years, into which major eras would you divide it, and what was the core question of each?
Chip History: The Long Sunset Under Monopoly
Liao Heng: That’s a long topic. But before we get into it, let me first describe how this came about. In the past, HiSilicon (海思) very rarely appeared before the media, so I feel fortunate to have received this invitation. I did think it over at the time — there were probably three reasons that led me to make up my mind to accept this interview. First, having such a chance to look back on this engineer’s story might leave a bit of inspiration for other people in this world.
The second reason: about half a year or more ago, at Tsinghua, at the invitation of Professor Lu, I taught one session of the Computer Organization course in the computer science department. After teaching that class I was left with a deep sense of frustration — and Professor Lu, who teaches the course, also felt deeply helpless when we spoke afterwards. In this era, young students are far more willing to chase hot topics — AI algorithms, model training, even inference acceleration. Computer Organization is a hard course to study; the big assignment is to design a complete CPU, and everyone regards it as a “gate-of-hell” course. So it’s very hard to attract students to take an interest in hardware — computer processor hardware. On the one hand I felt extreme surprise; on the other, a kind of anxiety: if people aren’t even willing to study this anymore, will the industry lose its follow-on inflow of talent — could a gap open up in the human pipeline?
And of course the third: we feel that people’s understanding of this industry — even my own subordinates or colleagues — is lacking. They’re buried in hard work every day and really don’t have many opportunities to be shown a more complete view of the industry — what the full story actually is — so it’s easy to feel lost or to waver. So I thought: if there were such an opportunity — not just for industry peers, but for young students, or for my own colleagues — to present a relatively complete, decades-long perspective on the industry’s development and its underlying patterns, that would be a very good opportunity too. So these three opportunities — three origins, let’s say — taken together helped me make up my mind to accept this interview.
Now, the topic you just raised — the broad arc of semiconductors over the past thirty years. I started my undergraduate degree in 1987; then around ‘97 — I went to the United States in ‘96, and in ‘97 I joined a semiconductor company. And of course the field I studied — during graduate school and undergrad I actually worked on processor-related topics. If we summarize the industry over these past 30 years, it has really had some very large rises and falls. Looking at it longitudinally, I think there are two main threads interwoven with each other that form the through-line of chips, or hardware. One main thread is obviously the processor, right?
The CPU — from about ‘86, or ‘84, slightly earlier than the IBM PC. But from around ‘84–’85 you had the IBM PC, and then the CPU gradually moved from a desktop office terminal into enterprise IT servers, and then into the entire global infrastructure that followed — with digitization, humanity entered a digital world, right? In all of that, the CPU is clearly the most important main thread. And of course, on this thread, starting around 2005–2006, this AI wave is a new wave again — we’ll probably unpack that in more detail in a moment. So that’s one thread: the processor-centric thread.
Then there is a second, very important thread: the chips brought by communications infrastructure. Over these past 30 years, humanity went from an unconnected state — a physical world relying on physical connection — into a connected one: everybody connected, every device connected. That course of development is the development of the internet, and it represents another longitudinal thread. The world went from no internet, no broadband, no mobile phones, to every person and every device being connected together — and of course that required enormous quantities of chips and enormous infrastructure. And this infrastructure went from very slow dial-up over phone lines and modems, to broadband, to optical fiber — that’s the wired side — plus, behind it, the trunk lines between cities and between countries. That is what built up the entire internet’s infrastructure.
Then, with the emergence of mobile networks — and later things like Starlink — that is, wireless infrastructure, this amounted to the internet’s second wave. In fact we are right now in that second wave of the internet: the mobile internet. On one side it spawned massive infrastructure; at the same time, on the device side, the most representative product is the smartphone, which has basically become a companion that we humans use for perhaps six or seven hours or more every day — it has become a necessary part of our lives. And this thing is also a huge, massive pillar of semiconductors: the semiconductors it consumed at one point may have exceeded 60 to 70 percent — that is, if you count all the phones. So those two threads are the longitudinal threads.
But if we then look by points in time, there is also a horizontal thread. What we see, dividing things up by time, is that semiconductors went through a phase of extreme, feverish ascent and prosperity — and then rapidly entered a state of withering and decline, even a sunset industry. The most emblematic symbol of this “sunset industry”: before this AI wave, roughly before 2016, for a full ten years or more — even fifteen — Silicon Valley’s, America’s VCs made no investment whatsoever in chip startups, OK? Because everyone believed this was already a finished, sunset industry. And the most representative case was Broadcom, which invented a “Broadcom model”: acquire relatively mature, profitable semiconductor companies, merge and restructure them, and cut costs — that is, raise operating efficiency. That is entirely a sunset-industry harvesting model. Then after 2015–2016 came this AI boom, right? So now, you could say: the internet bubble, then the mobile internet as the second wave, and AI as the third wave of these most recent years.
This third wave — the enthusiasm it has stirred up, the volume of capital, and people’s expectations for the industry — may be even higher than the previous two waves. So we should think about a question: why, after that previous wave — after the internet, say — was there a 10-to-15-year sunset period, a dusk? Roughly 2005 to 2015 — really it had already begun after 2000, yes. But from 2005 to 2015, almost a full decade, VCs invested nothing, because nobody was optimistic; people felt chips were a hopeless industry. I think here — and what I’m saying may not be politically correct — I think it was precisely the extreme success of certain sectors that caused the withering at another layer. I don’t know — maybe that’s a bit counterintuitive, isn’t it, Teacher Xiaojun?
Zhang Xiaojun: Yes — when you told me that, I found it quite counterintuitive.
Liao Heng: The logic here is actually quite simple. Of course, China has now begun to move into an industrial model of its own, a kind of self-sustaining cycle. But taking America before that — if you rewind ten years, global tech was still basically led by the American model, leading the world’s currents. And the American model has one striking feature: when a fairly important, great invention or a new business appears, and it gets the backing of the capital markets, they can rapidly build a near-monopoly advantage. For example Google, or Meta — that is, Facebook — in social, Google in search, and before that Windows in terminal operating systems, plus Apple’s phones. In their respective fields, each of them — in perhaps just three or four years — could complete the establishment of a monopolistic advantage. That advantage is not only a technological advantage; more important are capital, infrastructure, and the number of users — in effect, people get used to one thing and won’t easily switch platforms, right? And this kind of monopoly, this monopolistic advantage, in fact has an extremely negative effect on innovation. The effect is this: if a customer is the only merchant in the world purchasing a certain class of product, then as its seller — the supplier — the supplier’s lot is miserable. Miserable — well, you can’t quite call it miserable; it’s just what the supply-demand relationship inherently determines.
When a customer accounts for the vast majority of purchase volume, it has enormous pricing power, and it will inevitably leave the supplier with nothing to earn — gross margins become very meager. Second, they come to dominate demand and the direction of technology. That dominance looks perfectly reasonable, but let me try a small example: if you listen entirely to your customer, there is no future. Because in chips, it usually takes at least three years from the main architecture or product-definition decisions to the product actually landing, in real users’ hands: the chip design cycle alone is a dozen-plus months, the manufacturing cycle now is maybe nine to ten months, then you still have to build it into a complete system, and then test. So three years is unavoidable — maybe two and a half if you’re fast, four if you’re slow. And so — people often lack that foresight — the calendars don’t line up: I live on a calendar set in 2030, while my customer lives in 2026. There’s a four-year gap in between. If I can only blindly obey my customer, then I will lag by four years. Therefore, if this so-called technical leadership — the power to define the future — is placed entirely in the hands of one dominant party with an overwhelming advantage, it will inevitably suppress new innovative forces. Because even if you have the right idea, you get no chance to put it into practice.
And this series of causes is what drove the traditional wave — that is, the pre-AI wave — of semiconductor companies, the vast majority of them, rapidly into a state of decline: first, operations were hard and there was no profit to be made; second, the things they had hopes for in the future rarely got a real chance to be built. Now of course, what I believe is that this latest wave — at least in the AI era, perhaps on the hardware side from around 2016–17, with AlexNet and the continual progress of these DNNs, roughly from 2016–17 — up to now still has not entered an unfortunate monopoly, or a single dominant pattern like that withering period I just described. And there is a very interesting phenomenon here: this field is extremely, extremely active, and its progress every month, every quarter, exceeds everyone’s imagination. Frankly, up to now — and of course there is also our China, an important participant in the world, a player that has stepped onto the field — you could say that, at least among the people we can reach, it is a sky full of brilliant stars; this field is intensely active. Nobody can say that Sam Altman, or Anthropic, will be the final winner. Because every day our peers — especially those super-smart, ambitious young people, perhaps in Wudaokou, perhaps in Hangzhou — keep bringing us surprises, preventing the monopoly from ever forming. And so more of this vitality keeps appearing at every single layer of this industry.
Zhang Xiaojun: The dominant players you’re talking about — you mean Google and those big guys, right?
Liao Heng: I think we in China also have plenty of dominant players. For example, if you buy things — shop online — you can identify who the dominant players are, right? If you say short video, you also know who those dominant players are.
Zhang Xiaojun: That is, the dominant players of the internet-application wave.
Liao Heng: Right. They represent enormous scale, enormous purchasing power, an enormous consuming population — and they also sustain massive infrastructure.
Zhang Xiaojun: I said it was counterintuitive because in my imagination, the development of upper-layer applications and of the underlying chips should be a mutually reinforcing relationship. I hadn’t expected that once monopolists form at the top, it would instead leave the chip industry in a period of withering.
Liao Heng: Then let me ask you one question — you’re a hand in this industry too — when you picture Google, what kind of company do you think it is?
Zhang Xiaojun: It’s a search company.
Liao Heng: It’s an advertising company. To this day, the vast majority of its revenue still comes from advertising tied to this broad search business, right? I think the tech industry has one striking feature: nearly all of them — these entities we can identify — have a first axe-stroke, a formidable opening act, and it’s always impressive. Typically it goes: they invented something unprecedented, and brought humanity some enormous, wonderful experience, or some indispensable capability that they enabled. But it is very, very difficult for a company to keep remaking itself. Say OK — after my first-edition super trick comes out and I’ve established a monopoly position, do I still have the ability to invent again — to reinvent, to innovate again?
This is the classic innovator’s dilemma: someone who has already built a huge advantage in a certain field — he, or a company, an entity — will certainly reap rich profits there. But does he have the ability to reinvent himself in another field, to create the next, even greater thing? There are indeed quite a few brilliant people who can keep breaking through their own DNA, their own limits, and keep bringing humanity surprises. But for the great majority of organizations, the capacity to remake and recreate themselves — to build an entirely new self — is usually limited. If you don’t believe it, run every company you know well through your mind, and you’ll find they each have an advantage somewhere, don’t they? But an advantage is usually precisely a disadvantage in a new field. So we see that today in the AI field, the companies with the biggest breakthroughs are in fact not the big names we knew — the previous generation’s big names. That’s right. Of course the big names still want to — and of course they have resource advantages — they want to do this and that, they want to work hard at reinventing, and we very much look forward to them reinventing, since after all they have their own advantages in talent, resources, organization, and execution. But what I want to say is: the world is wondrous in just this way — it is often not the richest child who becomes the most successful. For example, look at the classmates we’ve known: was it the classmates from the best-off families who ultimately achieved the most in their careers, or contributed the most to society? Usually not.
Zhang Xiaojun: You entered Tsinghua in 1987, and through the Youth Class (少年班) at that. What was your computer science department studying back then?
Liao Heng: I’d say the coursework probably wasn’t lagging that much yet. Growing up — from my teenage years — what I got my hands on was an Apple II compatible; around the time I started university, the IBM PC appeared, right? And we were still on floppy disks — a hard disk was a novelty. Musk hadn’t founded his companies yet.
Zhang Xiaojun: Had Musk even been born?
Liao Heng: He should have been born by then. Right, but it wasn’t yet his time to start companies. So in that era, in terms of the atmosphere at school, there were a few striking differences. First, the teaching faculty of that time was still rather old-school and didn’t much care about publications. The research the professors did was all about building the real machine — at the very least a prototype. You could not graduate by writing a paper; it was always: I designed such-and-such, and I must see it — not merely conceive the thing, but actually produce an ugly circuit board that works. That’s quite interesting — it represented a culture of research. At that time, because China’s own industry — the research and R&D capability of its companies — was very weak (and going back earlier than my time, the gap was even bigger), universities and research institutes in effect took on the role that corporate R&D plays today, because companies had no R&D. That includes China’s earlier “Two Bombs, One Satellite” pioneers — the computers they built and used, everything they used. It happened to be a transitional period. The research system that came after me — including how students were trained and graduated — would shift into a publication-driven mode, right? More like the American, Westernized model, where everyone treats getting papers into top conferences as the first priority — and that was very much in step with the rhythm of the times.
Because when I left China, a company like Huawei probably had maybe only a few thousand people, and what they were building was honestly still fairly rudimentary. That is, the system for organizing thousands upon thousands of people to develop something extremely complex in an organized way had not yet formed — it was still taking shape. But if you fast-forward to today, many Chinese companies, big and small, have development capabilities — developing products, knowing how to organize people and resources, how to build a relatively complete human-resource system to accomplish that kind of teamwork and organization work — that are already very strong, even world-leading. So universities no longer need to take on the kind of development-type work they did in the past.
But going back to the era I was just talking about — what was particularly interesting is that back then, besides going to class, we actually spent a huge amount of time knocking around in Zhongguancun (中关村). “Knocking around” meaning — it was actually very much like today’s startups: OK, you discover the world has lots of needs, and then we would try to develop a product, or at least a product prototype, to try to meet those needs.
Zhang Xiaojun: What did you make back then?
Liao Heng: We made countless failed things — laser printers, VCD players, electronic dictionaries — all ancient stuff. People today probably don’t even know what an electronic dictionary is, right?
Zhang Xiaojun: I know what that is.
Liao Heng: Right. But at that time, companies’ R&D capabilities were very weak, so university students essentially filled that gap, developing these things on behalf of those small companies.
Since I myself was a competition student, I was decent at programming — passable at software. At the time I found my best friend, named Wu Zhao (吴钊), also a Tsinghua classmate; we were close in age, less than a year apart. He was a hardware competition student. So for many years I really yearned — thinking, I want to be like him: I want to be able to build boards, design FPGAs, and I absolutely must figure out what is actually inside a processor — what is really going on inside the chip. Right? Because chips have now reached over a hundred billion transistors, an extremely complex system. So when I felt my software was decent, I would pour enormous effort into trying to pierce through that layer — that membrane, that boundary — as if it were the difference between the upper floor and the lower floor of a building.
Zhang Xiaojun: That was your yearning at the time?
Liao Heng: That was my yearning at the time. And so, in much of what came later, I had a very strong preference — I was especially eager to understand what the world on the other layer was like. So later, professionally, I became through and through — or at least primarily — someone doing hardware.
Zhang Xiaojun: So at university you leaned toward software, right?
Liao Heng: My starting point was a programming-competition background.
Zhang Xiaojun: How old were you when you entered the gifted youth class?
Liao Heng: The gifted youth class generally required you to be under 15; I was 14 at the time.
Zhang Xiaojun: Wow, so young.
Liao Heng: I don’t think that’s anything special. The vast majority of people get in through happenstance — various circumstances give them a particular kind of opportunity. Getting into that kind of school purely on your own ability is quite hard, so it’s not worth dwelling on.
Zhang Xiaojun: What you were just describing is crossing between different layers — we can probably come back to this topic later. On the way here you also showed me Jensen Huang’s five-layer cake, right?
Liao Heng: Well, in our minds it’s more like an eighteen-layer pagoda — eighteen layers of “confusion.” (A pun: 宝塔, “pagoda,” sounds like 糊涂, “muddle.”) We’ll focus on that later.
Zhang Xiaojun: Back to 1987. When I was studying chip history I found it very interesting, because TSMC (台积电) was founded in 1987, I believe — Morris Chang (张忠谋) was 56 that year. TSMC pioneered the foundry model, which later rewrote the global chip industry’s division of labor: before, it was vertically integrated, and afterwards design companies became one category and chip foundry companies another.
Liao Heng: Yes.
Zhang Xiaojun: You lived through that era.
Liao Heng: By the time I experienced it, TSMC was already fairly powerful. Or rather — although it had not yet formed today’s kind of monopoly, that overwhelming advantage — it was already quite strong.
Zhang Xiaojun: How did this model come about? What is its root cause? Why was it a better fit for that era?
Liao Heng: I think the most fundamental reason is economics — it’s a principle of economics. Because a wafer fab’s capital expenditure keeps growing, right? Each generation it increases at — I don’t know — certainly something close to a Moore’s-law rate. That is, a process node that used to have maybe 10 mask layers now has 100 mask layers; each machine that used to cost maybe $100,000 or $1 million now costs $10 million or $100 million. You can think of the foundry — wafer manufacturing — as a business with extremely large capital investment and extremely long R&D cycles. This corresponds to what I’ll describe later as floors five and six of the pagoda — the basement levels. As a business it is extremely high-risk, with terrible ROI, and an extremely long payback period on investment — and as process technology evolves, it gets worse and worse. So at that point, as a design company, you simply don’t have that kind of capital capacity. It’s like how nowadays every home uses air conditioning and electricity — you wouldn’t build your own power plant just because you need electricity, because the outlay for a power plant is too enormous. So it became a kind of — I think Morris Chang was the first to realize this: that everyone would want to publicly share this infrastructure. Of course, on Bilibili you can probably find many interviews with him where he looks back on this history. But he was the first to realize it, and to explicitly turn it into a business model — that is a great achievement, a major contribution to humanity.
Of course, this contribution — countless companies have been built on this model and found success. I can put it this way: without the fabless model, even a giant like Google would have been unlikely to do the TPU. Because building a fab of his own might take three years; developing the process technology, another three years; getting the fab running properly would also take a long time. And its scale probably couldn’t sustain the economies of scale that a leading-edge fab requires. But once this model was established, it naturally enabled everything that followed — in any case, my entire career has been under this model; until about five years ago, it was all under this model, right? The model has its rationale for existing, but it is not everything. AMD founder Jerry Sanders famously said that only those with wafer fabs are real men — “real men have fabs” — but in the end even he spun off his own chip manufacturing division. This model has its strengths, but that doesn’t mean everything in the future can only follow this model. Because the rationale we’re talking about rests on certain big premises, and once those premises change, you have no choice but to adapt.
Zhang Xiaojun: What was the big premise at the time?
Liao Heng: The big premise at the time was global liberalization — “the world is flat.” Everyone specialized in their own part of the division of labor: at every layer, I want to procure the best capability in the world. So at HiSilicon, before 2019, nearly all — 95%, even 99% — of our designs were built on the world’s best suppliers: the “downstairs,” that is — I would pick the best fifth and sixth floors to build my product on. But that big premise has clearly been completely overturned by competition between nations, or by the influence of various other geopolitical factors. That’s the first point. The second is that when something develops too rapidly, this layered division of labor may not be able to meet the demands. For example, the extreme memory shortage that has appeared in the past half year has driven memory prices up more than tenfold — DRAM has risen more than tenfold in price. Vertical division of labor, to some degree, cannot cope well with such rapid — excessively rapid — change. It’s like the stock market, right? When there’s an explosive crash — a Black Friday — nobody can react in time; or when there’s an extreme surge, your normal trading operations can’t keep up with the speed of the rise.
Zhang Xiaojun: When did you realize you would probably spend your whole career in the chip and semiconductor industry? Was it at Tsinghua, or after you went to America?
Liao Heng: At Tsinghua I already really wanted to do this. Because of what I mentioned earlier — that yearning for an understanding of the “downstairs,” right? And I think the upstairs is more comfortable: the view is good, you can see farther. For instance, if you build an application, as long as you have the right concept, within maybe half a year you can rack up tens of millions of daily active users, right? The vast majority of super apps emerged that way. But the downstairs represents a more patient kind of player — a long-distance runner, so to speak.
Zhang Xiaojun: Are you that kind of long-distance runner?
Liao Heng: I don’t know. Throughout my career, most of my colleagues have said: you’re not a chip engineer, you’re a wolf in sheep’s clothing wearing a chip engineer’s skin, because you’re not really doing chips — you do software. And then when I’m with software colleagues, they either say I belong downstairs with the chip people, or they’ll suddenly say I do algorithms. So I think this polyhedral, multifaceted character — actually, doing chips requires knowing those things; or put another way, you need to be able to spar with others in those other domains without losing.
Zhang Xiaojun: I’m curious — in 1996 you went from China to America, to do a postdoc. At that moment, what was your imagined picture of the future of technology? What was the prevailing mood?
Liao Heng: That mood is a bit complicated to describe. I held two thoughts at the time. The first was that I felt I had to go, OK? Otherwise — even now, many excellent candidates I interview, say people who have already done PhDs at top schools and proven their abilities are second to none, still feel — many of my “Genius Youth” (天才少年) subordinates say: I still have to make a trip to MIT, otherwise I’ll feel I’m lacking. I had the same mentality back then. Because America as the technology leader — from roughly 1945 to now, I’d say, all these years — people would have a kind of regret: if I’ve never seen it, I won’t know, and I’ll forever feel something is missing. That was the first thought.
The second thought was this: I felt that in that startup-like mode I described earlier, I had made a lot of miscellaneous things that looked like products, but nothing I made had ever — say, sold a million units, or had countless people raving “I’m using the thing you built” — nothing had truly reached the market. At the time I didn’t know what my problem was — I genuinely didn’t — but after working for a year or two, I quickly found the answer. When I went over, I still carried that puzzlement: why is what I make never good enough, or never reaching the market? Actually the answer is very simple. After a year and a half of working, I quickly understood what the missing piece was.
Even if you’re making an electric kettle, you must ensure that even if one component inside fails, it won’t keep dry-boiling — otherwise it will cause a fire, or even kill someone. If you make an electric kettle, it absolutely must pass safety testing: a bit of water splashed on the side must not electrocute anyone; if the relay on top fails, it must not dry-boil; it must fail safe, not fail dangerously, right? That kind of process — for a thing to actually be usable, it’s not enough to have a clever enough mind to imagine how to design an electric kettle, or how to design a SpaceX rocket; you also have to execute it, execute it reliably, and you need sufficient verification and testing, before it can finally meet the standard of going to market and being used by ordinary people. As a student I hadn’t recognized this at all — I thought, I have a clever idea, I build it. In fact all we could claim was reaching a prototype. And that prototype — you can’t imagine: take a phone like the one you use — it may take tens of thousands of people in R&D, over ten thousand person-years of investment per generation, to get it to the point where it doesn’t, under some circumstance — black-screen the moment it gets somewhere hot, shut down instantly at a ski resort in the Northeast, or fail to charge. There are countless things that require a rigorous process, a strict process. Only after going through such a process, such a cycle, can the product be polished to the point of being usable.
For example, our intelligent driving — it was also first incubated in our team, but that process took seven years: from building a prototype that was already driving all around our campus, to finally launching and selling to the first consumer. Getting it to land was a long road — it took a full seven years. So as a student I had absolutely no awareness of this, and very naively thought that being fairly clever, or being able to program, or able to build a board, meant I could make a product.
Another interesting, related topic: only after a year of working did I realize that if you want to organize — never mind ten thousand people — even just 20 people to build a product, you can’t have every one of them be an Olympiad competitor. You need ordinary, conscientious staff — people whose intelligence or skills have simply come through an average education and your hiring screening process — to be able to work effectively, and you need to combine them together. You can’t pick a genius for every slot, and if they were all geniuses that would create problems too.
Zhang Xiaojun: Right, with 20 geniuses you’d definitely never manage it.
Liao Heng: All kinds of conflicts would arise among them, right? That’s another issue related to companies, or teams. So everyone has their own function. This is actually very healthy for us — whether for a company or for a society. Otherwise, only the strong survive — the ones just a tiny bit better than others, or with a slight edge in one particular respect, would leave everyone else no room to survive. But the reality is exactly the opposite: you need large numbers of conscientious, responsible people to finally get things done.
Zhang Xiaojun: In 1996 you arrived in America, at Princeton for a postdoc. What was your first impression of America?
Liao Heng: Well first, I had already been to America before that — I had gone in 1987. From then to 1996, another nine years had passed. On that first trip, I thought America was a fantasy land. I was only 14 then, and everything seemed like a dream, because the gap between China and America was just so vast. How silly was I? After coming back, I told others — or told the woman who later became my wife — America is so wonderful, they even have air conditioning out on the streets. Later, when I tried to recall why I had held such a stupid notion — in fact, I must have visited a shopping mall in Palo Alto, near Stanford. At that time China had no shopping malls, right? Back then we only had farmers’ markets and the shops lining both sides of the street. Well, a shopping mall is basically a street with a glass roof built over it. So my mistaken impression at the time was that America was so advanced they even air-conditioned the streets.
But by 1996 I was of course an adult. The first thing was a degree of disappointment. Looking at the world as an adult rather than a child, I felt a deep sense of not belonging. Because the Princeton I went to is an American elite-style institution — its temperament is rather different from other universities. Later I would actually learn that Princeton has many excellent qualities worth learning from, but at the time I didn’t feel that. I only felt, profoundly: this is a place where America’s elite aristocracy lives, and I do not belong to this class — so I need to leave this place quickly.
Zhang Xiaojun: So you stayed only one year and left right away.
Liao Heng: Right.
Of course I had a second disappointment as well. I had originally thought that maybe I could do a postdoc and perhaps have a chance at a faculty position. But by then I already understood — understood fully — that if I wanted a faculty post, I would have to do a PhD all over again, because the Tsinghua brand wasn’t strong enough. Today the world may be different — in academic progress, including integration with the world, Tsinghua’s gap versus that era has narrowed dramatically, and in places it even leads. Back then there was still a lot of inferiority in my heart. I wouldn’t say it was pure inferiority — perhaps a mix of arrogance and inferiority.
Zhang Xiaojun: Did you think about returning to China at that time?
Liao Heng: I didn’t think about it then — not at that time. I felt that going back then would be — I can only say it was probably a mistaken perception: the belief that returning meant you couldn’t make it, couldn’t hang on there, right? I didn’t want that sense of defeat. But later, the facts proved that this perception was very foolish. There’s no helping it — when you’re young you make some stupid mistakes.
Zhang Xiaojun: That probably doesn’t rise to the level of a mistake.
Liao Heng: I think it was simply a mistake. If I look back at classmates around my age, or ten years younger, who pursued this profession in China — even the average, ordinary people I described earlier, conscientious, positive, willing, not “lying flat” — they have all done very well. Why? A person’s growth depends partly on themselves, but even more on the larger environment, right? If the whole environment is rising, you naturally rise with it — like standing on an escalator that is itself moving upward — that kind of collective progress. Whereas in America, from my visits in ‘87 and ‘96, I felt they were not progressing quickly — they were even regressing. But China over these past 30 years has progressed very fast. And of course I was fortunate to return to China for the second half, and to take part in a piece of it.
Zhang Xiaojun: When I was reading chip history, I found that 1997 also had a big event in the chip industry: in April 1997, Nvidia launched its third-generation NV3. Only after the failures of its first two generations of chips did it finally gain a firm footing. Did you notice this company at the time?
Liao Heng: I didn’t. I’d say they were nobody at that time. In 1997, in the history of chip development, there was actually a far more spectacular drama playing out — the main storyline was not what you just described; that was the peripheral of peripherals, it wasn’t important. The main drama being staged at the time was actually an extremely important juncture for the CPU. Because at Princeton, I shared an office with some fellow students — and for those working on processors then, the company they yearned for most was one called DEC, Digital Equipment Corporation. DEC may have vanished from history by now, but at the time it was second only to IBM. DEC made minicomputers and mid-range machines, and was also in Boston. This was DEC’s final period. Digital had a flagship program at the time called Alpha — the Alpha processor. That processor represented the highest level of CPU design of the day, so everyone doing processor research wanted to join that team, to have the privilege of taking part — just as today everyone wants to work on Nvidia’s Rubin or Feynman, to be part of that program, which represents the highest level in the world right now.
Let me rewind this story a little — it actually relates to what I said earlier about how, once a monopoly forms, it suppresses innovation. The CPU started out single-core and single-issue — that is, executing one instruction per cycle, or even taking many cycles per instruction. Then gradually came the reduced instruction set processor (RISC), including its inventor, Professor David Patterson, who is still very active today, and Professor John Hennessy, who once served as president of Stanford. They led a processor-simplification movement: make every instruction very simple, so one instruction can execute per cycle. Then in the 1980s — by ‘87, perhaps starting from ‘84 or ‘85, and on through ‘87, ‘89, ‘90 — the era of multiple issue arrived, meaning multiple instructions could execute in a single cycle; multi-core had not yet appeared. And at that time there was an extremely important competition — probably still the most critical one in processors to this day: if a processor can execute multiple instructions per cycle, how do I schedule them? How do I generate a piece of software, how do I compile it into code, such that it executes several instructions every cycle? That way it runs faster. Because, as mentioned, at that time Windows — or rather DOS plus Windows — was already, on the desktop — on the CPU side a monopoly had taken shape, but in the server field it was still a hundred flowers blooming.
The hottest company back then was Sun Microsystems, plus SGI and a whole series of others — Digital was still around too — and the operating system was UNIX, in its various flavors. That was the most important golden age of processor development, and it lasted about ten years. The main battle of technical routes at the time: on one side, statically scheduled processors, called VLIW, where the compiler arranges the instructions in advance and, at execution time, each instruction stands for multiple instructions executing simultaneously; on the other, the superscalar approach, which is fully dynamic — that is, nobody pre-arranges which instruction executes in which cycle; the processor itself fetches the instructions, places them in a scheduling window, and schedules them dynamically. Why do I bring this up specifically? Because this important topic will come up later in our discussion of AI processors: all of today’s AI processors are statically scheduled — including the world-leading Nvidia GPUs. Everyone is in fact still walking the VLIW road, not the superscalar road.
That debate went on for ten years. Because software and IT systems had not formed a monopoly, besides Sun and SGI there were also HP, IBM, and Digital — a contest of many powers, like China’s Spring and Autumn and Warring States period: every school of thought was present, and everyone was actively experimenting, trying to prove their own doctrine right. But this story — let me fast-forward — ended in an epic disaster; it all came crashing down. First Digital went bankrupt; when it sold off its assets, the Alpha processor was acquired by Intel, probably for very little money. Digital’s biggest asset was a search engine called AltaVista — a forerunner that came about six or seven years before Google. AltaVista essentially pioneered the search engine. It just didn’t know how to make money at the time: it fetched a lot of money when sold off as an asset, but nobody yet knew how to turn a search engine into a business that earns money; technically it was a pioneer. Later, of course, very clever people invented better search-engine algorithms, and Eric Schmidt found the business model, so Google became a giant and AltaVista vanished into history. But what I want to say is that this debate still carries very deep lessons for us even today — many lessons to learn from. And as we keep moving forward with the evolution of AI processors, we may very well draw useful directions from this history.
Zhang Xiaojun: What was your own view at the time, in that debate?
Liao Heng: My view at the time was hardly worth mentioning — because what I was doing at Princeton was also VLIW. At that point I had never built a chip, never done a full-scale chip that genuinely had to tape out and go into manufacturing. So as a so-called PhD student I had no judgment — very naive. To have judgment, you need to know things, you need cognition of reality: first know what the situation is, be aware of your surroundings — OK, it’s hot today so I wear short sleeves; when it’s cold… You have to perceive the real world. That perception of reality, back then, I completely lacked. Later, after going through so much—
Zhang Xiaojun: After feeling that Princeton wasn’t quite the right fit, you chose to go into industry.
Liao Heng: I basically just went and found a job — whoever wanted me, I went. So I didn’t have many choices either. We all belonged to that cohort — I brought a thousand US dollars over, which probably already counted as being a rich man; plenty of classmates dared to go to America with a hundred dollars, some without even enough money in their pockets to pay the taxi fare from the airport to the school. So we had a fairly strong dose of reality back then: the feeling was, first of all, survive.
Zhang Xiaojun: In 1997 you joined PMC-Sierra. It had also just gone through a restructuring: a Canadian company and a Silicon Valley company had just merged, and you happened to join right at the point of their merger.
Liao Heng: Right. It wasn’t my active choice at first — they chose me. But actually it was a small company; at its very peak it probably never exceeded 2,000 people. What does that mean? The department I belong to today may have 5,000 — 4,000-plus people, with revenue easily 100 times theirs. So PMC was probably a small entity. But it lived through a magnificent, surf-riding industry pattern of the kind I described earlier. It happened to be standing right on the first wave of the internet, because its business was transport networks — several of the key devices in the internet’s backbone infrastructure, metro networks, including the long-distance transport networks between cities. That was exactly what it made. That wave was really just like the capital-market frenzy we see today, right? At one point PMC’s market cap became perhaps the No. 1 — at least the No. 2 — highest in all of Canada; it could easily have bought the Bank of Montreal. But that illusion vanished very quickly — a few good years, a few bad years. I arrived in 1997, and by 2000 it burst.
Zhang Xiaojun: Burst in 2000.
Liao Heng: In 2000 — you may not know; you look very young, maybe you were a child then, or not even born yet — but in 2000 there was a fantastic… anyway, it was the first bubble-and-burst cycle I lived through. The internet had become the world’s number-one hot topic; people believed everything in the world — that the internet would change everything. So on one front, Microsoft was in a fight to the death over the browser of the day, called Netscape: Microsoft built IE and was determined to kill Netscape — over the entry point, the portal. Just like now, aren’t all the internet applications fighting over the user entry point? It was the same then. At the application layer, people believed the browser was the first entry point, because everything had to go in through the browser, so whoever controlled the browser had the advantage. And at the infrastructure level, people believed everything in the world would pass through the internet, so the whole world was frantically, extravagantly overbuilding backbone and access networks, so that every city would have internet coverage, and the fiber between cities had to be laid with plenty of bandwidth. So PMC’s share price back then shot up just like Nvidia’s or Cambricon’s today — very, very high, right? But very quickly you discover — of course I’m not saying these present-day companies are the same — there is one biggest economic measurement in all this: whoever invests has to be able to earn it back. The internet bubble lasted about three years, and then people rapidly found that those investments… For example, the fiber America uses even today may still be what was laid in 2000, because so much was suddenly deployed and then — oh no — nobody’s buying, nobody’s leasing; you find it cannot be monetized. I think this problem does not necessarily exist in China’s AI today, but in the internet of that era, in America, it genuinely existed.
Zhang Xiaojun: Right — that was when you went through an internet bubble burst. But later the internet companies actually rose again, didn’t they? Around 2005 the internet application companies took off.
Liao Heng: Yes, the application companies rose again — but the people doing infrastructure never came back up.
Zhang Xiaojun: Then why didn’t you ever think of changing jobs, jumping ship? Was it that you had come to know people there — or, to put it another way, was it a question of them abandoning you versus you abandoning them, and you weren’t willing to abandon them?
Liao Heng: I never interrogated myself quite that explicitly. I just felt I still wanted to find another product — something that would let these chip-making colleagues, or myself, have the chance to build another product and bring in new revenue.
Zhang Xiaojun: And did you find one?
Liao Heng: Later we entered the IT industry — we started doing hard-disk, i.e. storage-related things. And the customers we served switched from telecom-side companies like Cisco, Ericsson, and Alcatel to IT-industry companies like HP, IBM, Dell, EMC, and Network Appliance. So I got the chance to learn: OK, this is what the telecom industry is like, and this is what the IT industry is like.
[1:00:00]
IT was once a very profitable industry. The equipment sold by companies like EMC and IBM carried very high margins. And the traditional IT OEMs — enterprises like HP — began to wither precisely after the hyperscalers appeared. Why? Because the so-called hyperscalers — those super-large internet companies — all need enormous infrastructure. In the global market today, I reckon that for purchases of servers, or of AI infrastructure, these big players — our term is OTT, over-the-top; the foreign term is hyperscaler, or call them super-large internet enterprises — should account for more than half of the global total, even 60–70 percent, even 80 percent. So when enterprises with that much purchasing power appear, they put enormous pressure on the brand-name OEMs. The old brand-name OEMs’ advantage was that they could make fairly high-quality products, but the customers they served were first the Fortune 500, then spreading out to the Fortune 1000 and 2000 — or, say, the tens of millions, hundreds of millions of giant, large, medium, and small enterprises worldwide. So an OEM needed a broad-coverage capability.
Right — say you want to sell servers into China: you probably need a service point in every county town, because that county might have a thousand enterprises as your customers, each buying only two or three physical servers; but when something breaks, you have to have someone make a house call. The hyperscalers are fundamentally different from that profile. First, there are very few of them — worldwide, maybe [unclear] — in the US perhaps the “seven fairies” (a colloquialism for the handful of US giants, akin to the “Magnificent Seven”) or however many; in China likewise roughly seven or ten, right? They are extremely concentrated. Second, their geography is concentrated too: all the machines sit in gigantic data centers, and you don’t need to set up a repair shop in some county town in Shaanxi, right? Third, when their purchase volume is 70 percent of what you produce, their pricing power is very strong. That is the other industry-changing driver I spoke of just now.
Zhang Xiaojun: And in that final year — the year you left PMC and joined Huawei, 2016 — the company itself was sold, for 2.5 billion US dollars. You happened to live through its entire cycle.
Liao Heng: That cycle was in fact precisely the sunset I described earlier. As I was saying, by around that time it more or less represented the entire American tech sector’s complete disappointment in, and abandonment of, semiconductors. First came Wall Street’s abandonment: their PE (price-earnings) ratios were maybe just three — on average about 3 to 5. What does that mean? Your profit this year is one billion dollars, and your market cap is only three billion. For reference: look at China’s AI chip companies today — their price-earnings ratios can reach numbers like six thousand or eight thousand. So first, the capital markets completely abandoned the semiconductor industry, because it had nothing new and no reason left to make anyone feel excited, right?
So what emerged at the time was the buyout model. Hock Tan — a genius of a leader, a business leader — invented swallowing the big with the small: first borrow a sum of money from private equity, then buy a company ten times your own size, buy it in at a low price, then restructure it — delete all the duplicated departments and personnel, cut the products with poor margins — and the financial statements turn beautiful. Operating that way — it is really restructuring-by-M&A, something that belongs to the late stage of an industry; effectively, certain people were liquidating the industry. Of course, the Nvidia you mentioned earlier belongs to the other category — injecting a brand-new hope into this industry — which is why we practitioners all deeply respect the contribution Nvidia has made.
Zhang Xiaojun: That sunset was truly long. From 2000 — three years after you joined the company — the dozen-plus years that followed were one long sunset period. What did it feel like to be inside it? I imagine anyone else might have left long before.
Liao Heng: I think the first thing was fear. Every day, colleagues would gather — over lunch, or while heating their lunchboxes in the microwave — and everyone would be discussing when the next round of layoffs would come and who it would be. But for me it was a — well, my wife also gave me a great deal of support at the time, or rather talked me out of giving up.
Because I think everything has two sides. First, it gave me ample opportunity to think and to observe: to think about technology — what kind of technology, what kind of product definition, has a chance to monetize; or to draw up a plan that contains the technology, the business, and a vision for the industry. And I had so much experience of failure, countless attempts. Also, although I was never laid off — right up to the last day — it gave me very much the role of an observer. For instance, I paid many visits to the IT giants — IBM, HP — dealing with them over the long run, going to Texas once or twice every month, and to IBM in North Carolina. I learned a great deal that way, including from the white-haired EMC engineers in Boston: from them I learned what this industry really was about, right — gaining more of a historical perspective. And I saw many failures, and came to know what kinds of things instantly trip the alarm — telling you this thing definitely will not work.
Zhang Xiaojun: Across that long sunset, what do you think were your biggest learnings? Looking back today, that stretch of experience may have been precious — because without it, the later Huawei chapter of your story might never have happened.
Liao Heng: What I learned — when I went to interview at HiSilicon (海思), the leader interviewing me asked: what can you bring? Because Huawei at the time wanted to hire “mingbai ren” — people who get it: they wanted to enter a certain field, but that field was new to them, so they needed to hire someone who had already worked in it for quite a while; that is what “someone who gets it” means. And I said: I don’t have many experiences of success — I have far too many experiences of failure, so I can tell you what won’t work. Of course I also often “kill the innocent”, right — my judgment is not always correct. But I have an especially large stock of knowing how things fail, and so the task is to work hard to steer around those failure factors.
Zhang Xiaojun: Could you name a few? Say three — failure factors, things you could tell at a glance would not work.
Liao Heng: For example, at the very beginning we set out to do ARM servers: we wanted to build a CPU to replace the x86 processor. My gut reaction — number one: failure. That was, at the time, a number-one product of what is now our Turing business unit. Back then I told the leadership every single time: whatever you do, don’t do this, because there is no market for it — or rather, no customer demand. Though look, I should stay humble here and admit my mistake: my judgment then was completely wrong. Because at the time CPU supply was plentiful, and as for the x86 ecosystem — AMD was not yet especially prominent, but everyone was used to Intel’s ecosystem, right? To change people’s habits you typically need a several-fold advantage: either you are cheap — a third of their price — or your performance is three times better than theirs. And at the time we possessed no such value advantage — that was my logic. But that logic, as I say, was wrong; the facts later proved I was totally wrong.
So my advice to everyone at the time was: we absolutely must find where the customers are. The only possible customers I could think of — because China’s cloud, China’s public cloud, did not yet have much scale; the cloud companies had all been founded, but they had not yet formed a good positive cycle — so I said the only opportunity was in America; we must immediately send people to Microsoft and Amazon. So we rushed straight off to Seattle. But unfortunately — in fact the action itself was not wrong; we were nearly on the verge of getting in, of having the chance to deploy this ARM CPU on Microsoft’s or Amazon’s cloud. That was 2016, 2017 — it started arriving around 2017; just as we had the opportunity to get into the large-scale mainstream cloud service providers as a CPU supplier, the US–China tech war broke out, and that story was terminated.
Zhang Xiaojun: That was still the previous era.
Liao Heng: Yes, that was that era. But after the tech war, when this story came back around, today we see that ARM servers really have — there are even predictions that within two or three years they may overtake x86 in the cloud, in the so-called data center, right? Because all sorts of people at these enterprises kept pushing — the hyperscalers even built ARM CPUs of their own. In China, too, we deploy on the order of a million-plus every year. And the rationale is not a pure value rationale — not purely “I have a threefold competitive advantage” — but involves many macro political factors. For example, China’s critical infrastructure can no longer run the risk of being penetrated through IT the way Iran’s power stations and nuclear infrastructure were, right? We must have something of our own. And our products have also improved very fast: no threefold advantage, admittedly, but no worse than the alternative — they have reached the point of being entirely usable. So society as a whole was fairly ready, and the product is also quite mature. That process took another ten years — but how many decades does a person get in one life? So I say my assertion back then was wrong — and yet, within that short time window, it was also not wrong. That is the first one.
Zhang Xiaojun: And the other two?
Liao Heng: The other two — well, as I said, it is about knowing what will not work and what will, right, and basically everything I can tell you is an example of my being wrong. At the time we wanted to build the Ascend (昇腾) chip, and the leadership said we absolutely must build a flagship product capable of training, capable of forming large clusters. I went and tried to persuade the leadership: whatever you do, don’t build this — it won’t sell. Because I had not yet seen the day when China’s sky would fill with brilliant stars — in 2016 I genuinely had not seen it. All the important algorithms and inventions were in America, right? Chinese researchers had not yet distinguished themselves — today it may just be Chinese in America competing with Chinese in China, but back then it was nowhere near so obvious. So I felt that a training-oriented chip — and Chinese internet companies had not yet begun training models themselves — I felt that product could never sell, and that we should build a small Ascend instead, right? When we first released the Ascend architecture, we hoped it would cover everything from one dollar to a hundred million dollars — six or eight orders of magnitude of broad coverage. And we did in fact achieve that: the smallest product sits inside our earbuds — those clip earbuds — and the biggest product is today’s training clusters of several hundred thousand cards, right? The spread we actually achieved is probably even more than eight orders of magnitude — whether in power consumption or in selling price. But at the time I was deeply pessimistic about the data-center product, and very bullish on the edge products. As the facts show, looking at it today, the data center—
I kept wanting to tell everyone: come on, let’s be realistic first and build something we can actually sell. But the facts proved me wrong yet again, right? So as I said: even when you hold that many prior assumptions (先验的认知), they are not reliable. And yet they still serve some purpose.
Zhang Xiaojun: Look, the several mistakes you just described were really all choices made on the basis of a kind of pessimistic expectation. Is that actually connected to the decline of the American chip industry that you lived through?
Liao Heng: Maybe there is. Maybe that experience trained me into a rather pessimistic habit — of first seeing the side of things that won’t work — and then working hard to escape that expectation of certain death, so to speak.
Zhang Xiaojun: And the examples you just gave all took place before the tech war.
Liao Heng: Yes.
Zhang Xiaojun: Including our autonomous driving — you mentioned just now the seven years, that we started doing autonomous-driving chips in 2019 — all before the tech war, when the external environment was still relatively good. So really, the choices you were making then were all small choices; you didn’t dare to bet big.
Liao Heng: At that time, I have to admit, some of the main leaders at Huawei — including the leadership of HiSilicon, or certain leaders at group level — had a much bigger frame of vision than I did. I was like a loach from a small pond suddenly thrown into this vast ocean, so I hadn’t yet learned that larger perspective. But after 2019, when we were placed in such a difficult situation, we had no choice but to learn to look at problems from a bigger vantage point. That is — take the examples I gave just now: at the micro level they all look fine, but at the macro level — I said my judgment was wrong — I was wrong precisely at the macro level. I didn’t see the bigger picture—
Zhang Xiaojun: Right.
Liao Heng: —didn’t see the expectations for the future, what the world would become. Really this is a limitation you fall into more easily when your experience is shallower, or when you’ve been a frog at the bottom of a well in a small pond for too long. Or, put another way, the younger you are the more easily you make this kind of mistake. I’m now in my fifties; back then I was only in my forties.
Moore’s Law
Zhang Xiaojun: I have a technical question I’d like to ask: Moore’s law in the chip industry — at what point did people feel it started to slow down? And what are the reasons behind that?
Liao Heng: First, Moore’s law has been slowing regardless of whether there’s a tech war — that’s an objective fact. Moore’s law has three aspects: its economics, its performance, and its energy efficiency — what we call PPA. As the area shrinks smaller and smaller, the cost averaged over each transistor comes down. That is what made electronics the only product in the world that is anti-inflationary — that is the greatest contribution Moore’s law made, right? If you buy anything else, what you bought ten years ago was certainly cheaper than it is now. Only with electronics is it the case that a ten-year-old electronic product today has no value at all, except perhaps some collector’s value. That is the economic improvement Moore’s law delivered: anti-inflation.
The second is performance: it used to be thought that the smaller the transistor, the faster it switches, so performance improved. The third is energy efficiency — energy cost: flipping a smaller bucket consumes less energy than flipping a big bucket. But at roughly 7 nanometers — or even as early as 16 nanometers — the economics of Moore’s law had already stalled: from then on, the cost per transistor not only stopped getting cheaper, it got more expensive. So the economic benefit was gone. The performance benefit is very small — yes, it is still slowly improving, but not by large margins. Where Moore’s law still genuinely delivers a benefit today is in energy cost: because the bucket has gotten smaller, the energy consumed each time you dump out that bucket of water is lower — that is, the picojoule- or femtojoule-level energy cost per switching event gets smaller.
Actually, this topic gets at something very fundamental. I can put it to you this way: the gap between China’s semiconductors and TSMC’s semiconductors — the main gap — is precisely in this energy cost. For example, if you ask me to supply the compute of a hundred thousand GPUs, I can supply it easily; it’s just that compared with the world’s number-one product, my power consumption is somewhat higher. But we have abundant energy.
Zhang Xiaojun: Right.
Liao Heng: So for us, for China, that is not necessarily a — well, first, this answer already has many facets. The first facet: nobody needs to worry that if we were completely cut off, China would have no compute. The answer to that is clearly no — we absolutely do have compute. The second question: am I more expensive? Not necessarily. The third: do I consume more electricity? Yes, I consume more electricity — there’s no way around it.
Zhang Xiaojun: But as you just said, China has energy, right?
Liao Heng: Our electricity supply is perhaps three times that of the United States, and we have a large amount of spare energy. Electricity for data centers in China may cost one-fifth or one-quarter of what it costs in the US; or compared with Singapore and other places around the world, on average our energy cost is about a quarter of theirs — roughly that level, a quarter or less. So first of all, we do not want to use this as the argument to persuade everyone to adopt something that consumes more power. All I can say is that as a floor — a fallback — this is not something to worry about too much, and China’s energy cost and energy supply are completely fine.
Zhang Xiaojun: So Moore’s law is still going?
Liao Heng: Moore’s law is of course still progressing. But to describe it more precisely, I think we can no longer use nanometers, a length scale — I think we should describe it by the number of atoms. Because an angstrom is one-tenth of a nanometer — that is, if you look at it on the angstrom, atomic scale, it has already reached the atomic scale. Therefore we know Moore’s law must have an end one day. And now, within the dimensions of a transistor, the smallest feature may already be at the angstrom level — at that level you’re counting in numbers of atoms, right? Going forward we might as well measure this thing in how many atoms, because in the end you can’t shrink below a single atom, can you? It becomes a matter of there being one or none — that is where its ultimate endpoint lies. But I think for now we should still hope Moore’s law can continue. Its economics, however, will certainly get worse — it will become more and more expensive.
And a second point, touching on what I described just now: if everyone has grown used to measuring how advanced a transistor or a semiconductor is with one particular yardstick, why not switch yardsticks? Ask instead whether a thing’s operation time can get faster — switch from a spatial scale to a time scale. Because what we actually need is the time scale: I need to do more operations per second. And then, can we switch to an energy scale: each time I flip a transistor, how much energy do I consume — at the picojoule level? You’ll find that when you switch scales, I think an essential difference emerges. For instance, when you evaluate a person — is it by their looks, or by how smart they are, or by whether they score well in math or in Chinese? Change the subject, and the same person may get a different exam score, right? So which subject you test becomes very important. This Tau Scaling Law (Tau定律) we’re talking about — Huawei’s time-scale counterpart to Moore’s law, unveiled in late May 2026 — tests the time scale: it says I just want the circuit to be faster. We could also switch to another one — Carnot-something, Joule’s law — saying what I test is: doing the same task while consuming less energy.
Zhang Xiaojun: A different exam question.
Liao Heng: A different question — if we change the question, perhaps the answer changes too, doesn’t it?
Actually, although these answers differ, they have strong factors linking them to one another. Any engineering guy has basically studied basic electrical principles: what governs the switching speed of a circuit is capacitance times resistance. Under Moore’s law there are two factors at play. Resistance is certain to get worse: the thinner a wire, the higher its resistance must be; and as resistance rises, the circuit slows down. The second factor, capacitance: the closer two plates are to each other, the larger the capacitance — so in that respect it gets worse; but as things shrink — as the plates themselves get smaller — the capacitance gets smaller. I can tell you why, past 16 nanometers or 7 nanometers, there has been little progress in speed: the essential reason is that resistance got worse, and capacitance did not improve significantly either. One factor makes the plates smaller, so capacitance should fall; but they also move closer together, which makes it rise again — the two factors cancel each other out. In fact, all semiconductor progress has come from capacitance shrinking, while resistance keeps growing — that is the essence of Moore’s law.
So we see that as it scales down, it doesn’t get faster — mainly because capacitance only shrinks slightly; with two counterbalancing factors, there’s no significant speedup. But because the whole thing has gotten smaller, fewer electrons have to be released, so there is still a gain in energy efficiency — just no gain in speed. As for cost: the cost of fabricating something this minuscule has certainly gone up. Those are the difficulties Moore’s law faces, as I described. But these difficulties don’t really matter, in a sense, because the whole world is counting on it — the scale of this industry is enormous — so everyone should keep pushing. Still, as I said just now: if we change the exam question, might the answer be different?
Let me give you a small example. If we want to reduce resistance, I should make the wire thicker, right? But if I want to reduce capacitance, I should minimize the overlapping area of the two electrodes. Now think about it: if two wires run parallel, their overlap area is large; if I make them orthogonal, the overlap is only that crossing point where they intersect, correct? So a structural change brings a gain. Plain scaling shrinks the original pattern: if I refuse to change the structure, that’s one answer; but if I say I can’t shrink any further and instead I turn two parallel wires into perpendicular ones — won’t the speed get much faster? The answer is: it certainly will.
This reminds me — you may have put a question to me before, saying one should ask a “cold question,” one whose answer is not so obvious. So let me give you this cold question — it actually speaks to the difference between Moore’s law and the Tau Scaling Law (Tau定律). Take a bold guess: in human history, how many patents have there been for mousetraps? Just guess.
Zhang Xiaojun: Three hundred.
Liao Heng: I don’t have the answer either — you could ask Doubao, or go look it up yourself on a patent website. I believe around 1900, the head of the US Patent Office submitted a proposal to Congress — wrote a letter — saying the US Patent Office should be abolished: the organization no longer needed to exist, because humanity was so clever that everything in the world that needed inventing had already been invented. One of the examples was that there were perhaps a thousand kinds of mousetraps — a thousand patents, of every variety, for exterminating that pest, the mouse — since humans basically all hate mice. So possibly by 1900 the US Patent Office already held over a thousand patents on how to kill mice.
Why is this an interesting question? Because back in 2020, I was making a serious effort to learn how lithography tools are built, and I found a textbook on electromechanical systems design. It left a very deep impression on me, because on the very first page of that textbook there was a diagram, and the diagram illustrated ten ways to kill a mouse. I can’t lay my hands on the book right now, but let me try to describe it for you. If you’re a chemical engineer, you’d say: to kill a mouse, I’ll concoct a poison, and then work through something the mouse loves to eat. Of course, killing mice always involves a lure — you need something the mouse loves, either cheese or some piece of meat, right? Put the poison on it, the mouse eats it and dies, correct? If I’m a mechanical engineer, I’ll build a trap — or a hundred kinds of traps: when the mouse comes for the bait, it trips the mechanism, and then I either snap it dead or shut it in a cage it can’t escape — in short, all sorts of methods, right? And if I’m an electrical engineer, I’ll rig up a high voltage: the moment the mouse passes my spot, it triggers a thousand volts and is killed instantly.
What I actually want to say is this: whether it’s Moore’s law or the Tau Scaling Law, it is simply human engineers solving problems with an Edison-style mindset. First: what is the problem? Second: have I found an effective way to solve it — to face the problem and find a good answer? Of course, if you’re making a product you also need economics — the solution has to be cheap enough. And fourth, the solution has to be repeatable, right? You can’t have quality failures — where this one unit solves the problem but the next one you build is unreliable, correct? All these factors stack on top of one another. And this is the big-premise issue I mentioned earlier: sitting here, if we start thinking about how I might reproduce TSMC’s or Nvidia’s capability — first, I don’t need to; second, however hard I think about it, I don’t have it in me to do so — my preconditions are simply different from theirs.
Zhang Xiaojun: You don’t have that context.
Liao Heng: Right. So a lot of what we’ve been saying has actually already drifted toward a metaphysical, or even philosophical, perspective.
Zhang Xiaojun: Just now we reviewed the chip history you lived through. For the second part, I’d like you to — imagine you’re a tour guide, taking us on a horizontal survey of this whole industry, the whole supply chain. Because you told me it’s an 18-layer pagoda — you don’t accept Jensen Huang’s five-layer cake; you think there are 18 layers. Take us on a tour of these 18 layers — a horizontal overview.
The 18-Layer Pagoda
Liao Heng: As we were describing — well, perhaps because our cultural backgrounds differ: Jensen Huang probably often eats French desserts, so he uses a cake as his analogy. As a Chinese person, when I picture a layered structure, the first thing that comes to mind is something like our Yingxian Wooden Pagoda (应县木塔) — right, the pagoda. But in substance it’s the same; we’re talking about the same thing, aren’t we?
That is to say — setting the pagoda itself aside for a moment — everyone has surely heard of the concept of co-design, so-called collaborative optimization: software-hardware co-optimization, or co-optimization between algorithms and chip infrastructure. The moment you discuss co-optimization, you’re already involving two layers, right? It means the classmates on the seventh floor and the classmates on the sixth floor know each other, know what each other’s difficulties are, work together, and make mutual concessions and accommodations: “Hey, your memory bandwidth is short here, so I’ll find a way in the algorithm to save some bandwidth; you have compute to spare, so I’ll spend — waste — compute to save bandwidth,” right?
Now let me point out a very intricate, very microscopic level of this. If you look at the Blackwell chip and today’s Ascend (昇腾), you’ll find a difference between us that perhaps nobody has ever described at this fine a grain. Let’s talk about Vector compute versus Cube compute. The Cube — they call it the Tensor Core, we call it the Cube — is the 3D matrix-computation unit we described: we name it Cube, they name it Tensor Core. Say their ratio is 32 to 1, and ours is 8 to 1. What does that represent? What important implications does it carry? I think the gap is perhaps a bit like this: imagine a family of four. When young, maybe they can only afford a 100-square-meter apartment. Later your career goes well, you’re a bit older, you’ve accumulated some wealth, and you buy a 200-square-meter luxury flat. Maybe five years later you get promoted again and buy a 400-square-meter villa. But aren’t you still the same family of four living there? You’ll find that when you live in the 400-square-meter villa, you probably pile up a lot of junk — say, delivery boxes, many never even opened, heaped in some corner nobody visits, right?
Zhang Xiaojun: Lots of redundancy.
Liao Heng: Lots of redundancy. What we’re getting at is: if you were suddenly told to move back into the 100-square-meter apartment, you could actually still live your life.
So 8-to-1 versus 32-to-1 means that when you go to use this compute, your model has to reckon with it: if my space is plentiful, I can waste it freely; if my space is scarce, I must find every way to use it fully — either build something “economical and practical,” or build storage shelving and clear out everything unneeded as much as possible. And you’ll find, if you look at the DeepSeek model, that they made some very deliberate, very conscious choices — I think their sense of direction is extremely clear. For instance, before 2025 they had already realized: I absolutely must economize on compute. Because of the scaling law, everyone believes more parameters bring stronger capability — but does every single parameter really have to be brute-force computed for the model to obtain that capability? Can I compute only one thirty-second of them? I think DeepSeek’s V1 — or V3, and R1 — amply proved the point: you don’t need to. That is sparse activation of parameters: if I just cleverly select one thirty-second of the parameters to participate in the computation, don’t I save thirty-two-fold on compute? Memory access is also cut thirty-two-fold. When people consciously confront this problem, they gain a 32x speedup.
Then of course, in 2025 to 2026, they did another thing with a very clear direction and goal: what about long sequences? As sequences get longer — a 4K sequence versus a 1M sequence differs by maybe 256x — must you pay a 256x computational cost to handle the long sequence? In fact, no. They did compression, sparsification: out of that 256-fold of tokens, I don’t need to compute every one; I select 1K, or 512, of them to compute, and I save that compute. But notice: such a selection process exacts a price in complexity. So you can see this model doing co-design very meticulously and very consciously.
Zhang Xiaojun: Right.
Liao Heng: They consciously said: we have only the 100-square-meter apartment, we still have to house a family of four, we still have to raise our child — and we want our child to be outstanding, no worse than Musk’s kids. But that requires pouring in more effort and taking on greater complexity.
In fact, that complexity comes precisely from the vector computation I mentioned. So from this angle we can say DeepSeek is extremely well matched to 8-to-1 compute: in order to pick out the items to send off for brute-force computation, it spends more vector computation on the selection. Whereas I believe Anthropic’s models — or OpenAI’s models — don’t need to do this, because the machines they buy come with 32x Cube compute. So let me put it another way: our current chip happens to be one-quarter of Blackwell’s spec, yet when running DeepSeek it seems to do fine.
Zhang Xiaojun: That’s co-design.
Liao Heng: That’s co-design — that is spanning two floors of the building: the algorithm folks realized where their model’s advantage ought to be built, right — that’s co-design. And this co-design has actually meant that with this kind of model, if I run inference on a DeepSeek V4 Pro, compared with the Claude — what is it, 4.8, or Fable — that I use every day, I don’t feel its token output is any slower, right? Of course we still need to respect what others do well — catch up where we should catch up, surpass where we should surpass, no? But — I can—
Whether we are the ones working on chips or on models, if you have this layer of consciousness, this understanding of the big premise, the direction of our efforts will differ enormously. I clearly have to give Liang Wenfeng (DeepSeek’s founder) a warm round of applause, because he consciously and proactively chose higher complexity — algorithms that are harder to make converge. With algorithms this complex, convergence can be difficult and all kinds of anomalies appear, and he worked hard to overcome those problems, because he had advance awareness. It is not that he waited until five years later, suddenly at a dead end with no compute at all — he had the foresight to choose to solve, ahead of time, a problem he anticipated. That is the value of someone with foresight: in effect he led the way, and others will keep pushing along that path.
Before 2019, we too were living in a happy world. HiSilicon (海思) back then was very hidden, very low-profile, but our scale was already very large — perhaps among the top one or two in the world. In terms of wafer purchase volume and product variety, we were already at least a top-3-scale semiconductor company in the world — it is just that all of its products served only Huawei’s own products. But what I mean by the “happy times” is that everything you procured was the world’s best, because you could afford to pay. For example with wafers, we would use TSMC’s most, most advanced wafers — even processes that nobody else in the world, including companies like Nvidia, would use yet. We were the first to use them, and if not first then second, because the position our products occupied could support that kind of spending, right? But once all of that was shattered, we had no choice but to face much harder things.
The second thing I have learned over these past years is that this, too, can be done. When you are pushed into a corner, you simply have to learn what is going on “downstairs”: how exactly to enable 7nm and 5nm process nodes, what efforts are required. I think perhaps the most valuable point is this: a problem that looks impossible, once you decompose it, becomes ten or a hundred more concrete problems; then you decompose those hundred problems one more layer, and it may become a thousand problems of physics, chemistry, and mathematics. This decomposition takes the macro question — say, today you ask: can you be as good as TSMC? The simple answer is no. But that answer does not help you, so what do you do? You start decomposing one layer down, and you find: if these particular places do not work, then which places do work, right? So an unsolvable problem becomes concrete, becomes tangible. Once it is tangible, out of 100 problems you can maybe solve 80, and you find you are apparently not doing so badly, right? Then you improve further on those eighty, and you find that on ten of them you are doing better than others — maybe you can play to your strengths and use certain advantages to compensate for your weaknesses.
Zhang Xiaojun: So what are the eighteen layers, exactly?
Liao Heng: The eighteen layers — first, the very top layer is applications: for example Doubao (豆包), or WeChat, or Ali’s Alipay. These are the application layer — the tip of the pyramid.
Zhang Xiaojun: The tip of the pyramid.
Liao Heng: Because it can monetize directly, and while it certainly has its technology, it is not necessarily an Einstein-level technical challenge, right — it is more about operations or business design. The bottom-most layer is very basic stuff: ores, mining, going to Nigeria to dig a mine — the bottom layers are essentially physics, chemistry, mines and the like, because a semiconductor is in essence a pile of sand plus certain rare metals. A bit further up is device design — say, whether you use a FinFET transistor or a GAA transistor, or what I mentioned earlier: if I can change the transistor design from two parallel lines to two perpendicular lines, does my speed not suddenly become ten times faster, because my capacitance shrinks a great deal? A bit further up again: how do you manufacture trillions of transistors on a single wafer, manufacture them reliably, so the chip can still work for five or ten years without failure? That is extremely hard.
But fortunately, the Chinese nation also has a great many very capable people — some Taiwanese, some mainlanders — and as a whole they have mastered these capabilities, right? So when they run into difficulty — constrained by machines: you do not have the most advanced lithography tool, you do not have the hundreds or even thousands of kinds of most advanced equipment — there are also countless people working hard to build that equipment, overcoming one obstacle after another. What I want to say is, this already brings us to roughly the fifth and sixth floors, right? The process node: do you have a reliable 5nm process? If you cannot do 3nm, what then? If you cannot shrink by another level, can you stack? What machines does stacking require? How do you stack these things reliably without creating heat-dissipation problems and power-delivery problems? These are the fifth- and sixth-floor problems downstairs.
Coming back to the seventh floor — that is roughly the level where chips sit: how best to exploit what is below, facing your own physical constraints, right? Because as we said, if our basic premise differs from others’, our process node is also different — where others need no stacking, I need stacking — you must face those physical constraints. And looking upward, we need to design the next chip’s architecture four years in advance so that software is relatively easy to program, right? So above the chip sit compilers, programming languages, the most sensible ways to partition for parallelism, reinforcement learning; then how to do KV cache, how to do model compression, how to do sparsification, how to invent a new data format so that algorithm people can use lower precision without harming model performance or preventing convergence. From the algorithm level upward, everyone is probably fairly familiar, right? Once you have a good slow-thinking model, or a model with agentic capability — how do you build a KV cache system, how do you build an agent framework, and finally construct a valuable application. Those layers are visible to the great majority of people. So above the seventh floor are the events everyone can see; below the seventh floor is the basement, where the crowd that ever looks is rather smaller.
Zhang Xiaojun: Which of these layers are the ones you know well?
Liao Heng: As I was saying — since I do not do any single layer well enough, basically all I can do is “muddle about”: roam between these layers, helping the colleagues who are more specialized at each layer to bridge the gap. Because it is very hard to be both deep and broad. It is like digging a hole: you can be a very sharp needle — stuck into a watermelon, you can pierce right through it; or you can be a knife and cut the watermelon in half. You cannot really be both a needle and a knife — you cannot be both sharp and broad. But I think this is exactly the point: we are actually not particularly short of experts at each layer — not just at Huawei, but looking across the whole world — because each layer is a profession, and once people have worked in a given role long enough, their understanding of it naturally keeps growing and they accumulate experience. What is relatively scarce is the ability to cross layers. As I said just now, maybe there are a hundred thousand AI algorithm researchers in the world, and maybe only Liang Wenfeng alone recognized deeply that he needed to break through that sparsification problem — to trade greater complexity for less compute. You see, that is a choice of direction; many people would not choose it. So this crossing of layers is the core of co-design. When a person can span those levels, they are like a thread stringing many pearls together, forming a necklace.
Zhang Xiaojun: Laying these 18 layers out, where does China have relative strengths, and where is it relatively weak?
Liao Heng: I think the strengths are obvious. First, the application layer is a clear strength — things like Alipay and WeChat Pay are just not as convenient anywhere else, right? These days, basically — I probably have not seen renminbi cash for several years now; I have not touched money, I have not touched cash in years. But if you travel anywhere else, you find you either still have to carry a credit card — otherwise you may not even be able to check in at your hotel — or you still have to carry some cash, worried that in Japan, say, you will not manage to buy a train ticket, right? So at the application layer we have a great advantage, and Chinese people are already very digitized, especially among consumers. And all the way down until you reach the chip, I think China’s strengths are at least enough to be no worse than anyone else’s. In algorithms it is even more a sky full of stars — I think perhaps 70% of the world’s outstanding algorithm people are Chinese. Why? Because a while ago I saw a Berkeley professor — a science-and-engineering professor at Berkeley — who wrote a piece saying that in his classes he now has to teach students the distributive law of multiplication: A times (B plus C) equals A times B plus A times C. So the gap is enormous. Chinese parents all take their children’s education very seriously; we have the imperial-examination (科举) tradition, the idea that “he who excels in learning becomes an official” (学而优则仕); and the upward channel for our people still exists — through your own effort, test into a good university, learn more, and you still have a chance, right? Therefore China’s talent supply is not lacking — indeed it is a far, overwhelming advantage. And of course our ability to monetize is not lacking either, because Chinese people pursue the good life — everyone wants to live better, earn more money, or achieve greater success, whatever the measure of success may be.
Zhang Xiaojun: Chips seem to sit right in the middle of this 18-layer pagoda — below it is the basement, above it the building; it occupies a middle zone.
Liao Heng: Yes.
Zhang Xiaojun: A connecting point.
Liao Heng: Yes.
Zhang Xiaojun: You came back to China in 2016 — between then and 2019, did your feeling about the whole chip industry also change a great deal, even though it was only three years?
Liao Heng: I think in 2016, I did not believe.
Zhang Xiaojun: Did not believe what?
Liao Heng: I did not believe that we, China, had this capability.
Zhang Xiaojun: Then why did you come back—
Liao Heng: I came back because of my wife. My wife said firmly that once we had a child, we absolutely must not let him grow up as an ABC (American-born Chinese), without an identity of his own, or trying to find his personal sense of self amid a kind of contradiction.
Zhang Xiaojun: How old was your child then?
Liao Heng: I came back as early as 2008, because my child was three.
Zhang Xiaojun: You came back in 2008—
Liao Heng: In 2008, although I was still employed at the same company—
Zhang Xiaojun: But you were physically back in China.
Liao Heng: I was already physically in China. Because I think my wife came to this realization earlier than I did; she said even more firmly: do not stay abroad, we must start over in China, especially considering the next generation. She was utterly determined to return to China, so I did too — at that time returning to China was a passive choice.
But by 2016, when I joined HiSilicon — I think perhaps the bigger shift in what you would call professional perspective came with that move. Before 2016, I completely did not believe. That is — please note, foreigners have a kind of arrogance, a blind self-confidence. The first element of this blind confidence is believing they are the best. Of course, once you have worked in that environment long enough, you become infected with this mistaken perception too — believing that a certain thing, only they can do. And “they” included myself — I was one of them, right? Although Chinese, having joined such an environment, I was part of that team — so you come to believe that for the harder things, only “we can, nobody else can.” But as for how things in this world actually work — as I said just now: take a very hard problem, and when you decompose it into 100 sub-problems, you find every one of them becomes more concrete and solvable. If someone is willing to solve those 100 problems, perhaps they will solve them better than you.
That is exactly the transformation I underwent starting in 2016. My first shift was this: things I had believed completely impossible for China’s industry at the time — in fact, our capability far exceeded my expectations. And I find this very interesting: I believe the great majority of Americans, American practitioners, still hold my 2016-era view. That gives us a huge opportunity. Because if you regard them as a competitor, and your opponent severely underestimates your capability, you hold an enormous advantage. They innately believe they can compete with you while living comfortably — clocking off at five every day, having a coffee, then going surfing. In reality, at that point, if you just push one notch harder, you can easily overtake them. That is my biggest realization: whether in 2016 or for many years before it, our Chinese counterparts — the effort they put in, their understanding of the problems — so long as their identification of the problem is at the same standard as everyone else’s, their ability to solve it far exceeds that of practitioners abroad, because they work harder. Put simply: you reap what you sow. But the precondition is that you need the right people to lead them to the right definition of the problem. Because if the problem is defined wrongly, all the effort is wasted.
Just now I gave the example of Liang Wenfeng. He defined two problems. First: model parameters must be sparsified — and he solved it. Second: sequence length — attention must be sparsified — and he solved that quite well too, right? You see, the definition of the problem itself — if we are allocating credit — already takes 80% of the credit; actually finding a concrete method to solve the problem once defined perhaps takes twenty percent. It requires talent, requires intelligence — maybe that intelligence comes from an intern — but the definition of the problem comes from the team’s leader, because the leader must steer how others should go about solving problems.
The History of Ascend and China’s Path
Zhang Xiaojun: Now we come to the very heart of this interview — I would really like to hear you tell the history of Ascend (昇腾), because these past few years your development has also been very low-profile. If you used three words to describe the past decade, 2016 to 2026, which words come to mind?
Liao Heng: I think it has just been a tribulation.
Zhang Xiaojun: A tribulation.
Liao Heng: Yes — or rather, a tribulation with hope in it.
Zhang Xiaojun: From the 910 to the 950, the environment you faced was actually very different. Standing at those design points, do you think the problems you defined for the 910 and the 950 were different?
Liao Heng: Yes.
Zhang Xiaojun: What were they, respectively?
Liao Heng: At the time of the 910, the basis we worked from was possessing the world’s most, most advanced logic process. But at that time, regarding the AI ecosystem, the development of models, and what their applications would actually be, we were fairly lost. In fact the whole world was somewhat lost then, right? Although I believe Nvidia was probably the earliest to sense that this area would explode, I believe that worldwide at that time — because model capabilities were not there yet, and no enormous scale had emerged — the confusion was more about how to guess which direction the future would take.
At that time we had already realized we needed to do model training; the inference market simply did not exist at all, right? So on edge devices — the earbuds and phones I mentioned earlier — we tried to find a path to monetization on the device side. That is, we thought data centers were for training, and elsewhere, trained models would be deployed onto all kinds of small devices. We did not foresee this kind of large language model inference — including things like AI coding; none of it appeared in our imagination. So at that time, it was merely that the conditions were very good, but we did not know where our own biggest future monetization opportunity lay.
[2:00:00]
Then by the 950, things had reversed, right? The conditions were extremely harsh: the opponent you faced was already Silicon Valley’s — the universe’s — biggest company, the one with the highest market capitalization and the strongest capabilities. So at this point, how could we survive, and then how could we serve — perhaps we harbored no over-lofty ambition of seizing the world’s number one position; it was simply this: step one, stay alive; step one, serve the customers who need us — satisfy that basic demand of being “customer-centric”.
Zhang Xiaojun: And what would have happened if you simply had not done it?
Liao Heng: We could also have not done it. What I want to say is that in reality, whoever this world loses, the Earth keeps turning just the same, right? It is only that, on the one hand, we still felt there was something to hope for — OK, we did not want to “lie flat”, so we still wanted to make an effort; and on the other hand, many industry peers probably had expectations of us, so we felt we did not want to give up so easily.
Zhang Xiaojun: What was the change in your state of mind? The difference when designing the 910 versus the 950 — the two generations of chips before and after the supply cutoff — that change in state of mind.
Liao Heng: I think putting it into words may sound a bit — what I want to say is: try to imagine, if you were Dong Cunrui (董存瑞) — the PLA war hero who died holding a demolition charge in place — about to hold up that satchel charge, right — we ordinary people cannot imagine what his state of mind was. Or a soldier at Shanggan Ridge (上甘岭) — what was his state of mind? Or — I once saw an interview from the Second World War with a Chinese soldier, also a story that appeared on one of Huawei’s promotional posters. A reporter asked him: what do you want once the war is over? He said: I don’t think about it. Because my parents are dead too… there is no point thinking about those things. So I would say the so-called change in state of mind no longer carries much meaning. It is simply: faced with a difficulty, you just want to solve it.
Zhang Xiaojun: You said that when you first joined Huawei in 2016 you did not believe either. When did you go from not believing to believing?
Liao Heng: That actually happened quite quickly. Because my disbelief was based on a foreigner’s arrogance — blind conceit — the belief that the capabilities they possessed could not possibly be possessed by anyone else. Reality educated me very quickly. After I joined HiSilicon (海思), I very rapidly came to understand that the colleagues around me, although different from me — for example, HiSilicon might newly tape out perhaps 100 chips a year. In any case, our President He (He Tingbo of HiSilicon) said at the Tau Scaling Law (Tau定律) launch event that over five years we had done 381 chips — that averages about 70 a year, right, 70 new tape-outs per year. But I have never once heard of a chip coming back “smoking” — that is, failing test and in the end proving impossible to mass-produce.
What does that show? Such a high hit rate shows that our colleagues — every single team — are highly qualified. Regardless of their origins, or whether their résumés look glamorous, in the domains they are responsible for they are all extremely responsible. As I said earlier about being an engineer — this is not about myself — it is something I did not grasp as a student, and grasped after a year of working: to be an engineer, you must first of all be dependable, and you must be responsible. That is the foundation of everything. And therefore, a hit rate like that would rank among the best in the world — and the outcome gave me a rapid education and transformation.
Zhang Xiaojun: I believe that no matter from what period you look back on Huawei’s history, the supply cutoff (断供) will count as a hugely significant event. When it happened, what was it like in your department? What was the reaction in that first moment?
Liao Heng: I think President He’s letter must still have been written before the cutoff. (Her internal letter at the time of the US action became famous.) But once the cutoff truly came, there was never again a chance to write that kind of letter. As with the example I gave earlier — you cannot imagine what Dong Cunrui felt. I can only say that everyone was in a different role, and their feelings were not all the same; I have no way to describe it.
I would just say, first, there were perhaps two sides to it. On one side, you needed to reassure the colleagues below you, tell them not to be afraid — because enormous fear arose in everyone’s hearts. Second — speaking for myself, and my memory is a bit blurry now — the deepest feeling I remember is that it ignited an enormous competitive drive in me, an enormous passion: this absolutely had to be solved. First, I had to know exactly what problems we had run into; then I had to break those problems down one by one and solve them one by one.
For an engineer, this was a godsend. Because under normal circumstances you never need to understand how a motion stage achieves one-nanometer precision, or how to measure something’s position. If you want to control a stage to move with one-nanometer precision, you also have to locate it during high-speed motion, know exactly where it has moved to; then you have to work out whether to measure it with a laser or with something else, right? These things are all interlocked, and you run into a huge number of interesting technical problems of the kind engineers need to solve.
Zhang Xiaojun: Did the team atmosphere change before and after?
Liao Heng: I think the team atmosphere probably did change quite a lot, but overall it was all right. Because precisely in the most difficult period, I think maybe only 10% — 5% — of the people around us quickly went looking for a new way out. But maybe eighty or ninety percent — especially the backbone, the most capable colleagues — I believe that to varying degrees, perhaps like me, each had their own problem-solving passion ignited. So many people — the great majority — chose not to overthink it: the problems, once broken down, were fascinating, so they just went and solved them.
That describes just that one phase. If we want to put it this way: every human heart has two sides, good and bad — the side that wants to survive and get more benefits, and the side with certain visions, or greater ambitions, a greater sense of mission; it is only a question of which side gains the upper hand. It’s like Gollum in The Lord of the Rings, right — that creature who on one hand wants to seize the Ring for himself, to gain greater power, but on the other hand feels that this thing is not quite right.
Zhang Xiaojun: The most difficult period — roughly from which year to which year?
Liao Heng: I’d say it was difficult at every stage — it’s just that the difficulties each stage faced were different.
Zhang Xiaojun: And the very lowest point?
Liao Heng: The lowest point was when there was completely nothing — that is, people felt that within maybe a few months we would be completely unable to manufacture chips, and the entire commercial loop would therefore be broken. But fortunately, those difficulties were essentially carried by a few exceptionally resilient leaders, who locked them tightly within a very small circle of their own. Those people bore the great majority of the pressure from these difficulties; everyone else remained more or less in the dark.
Zhang Xiaojun: And presumably there were also some landmark events?
Liao Heng: Perhaps there was no single event. But Huawei — take China’s 5G network, which is now everywhere — throughout that process we never once broke supply, and never caused construction of the entire communications network infrastructure to stall, or its schedule to slip noticeably. Behind that were the tireless efforts of countless people that made these things possible. I can say that perhaps upwards of five thousand, maybe as many as ten thousand boards were all redone — every item had components swapped out, and all of them rapidly reached mass-production capability again. That was the result of the joint effort of an extraordinarily large number of people.
It’s just that these things have no single marker — I mean, when something does not get cut off, what does that even count as, right? It’s what everyone takes for granted in daily life: no downpour today, no earthquake today. What I really want to say is that behind this outcome were very many people’s efforts, day and night.
Zhang Xiaojun: Why is Ascend (昇腾) called Ascend? And when was the name settled?
Liao Heng: I may not remember this very clearly anymore. It should have drawn on the Chinese phrase ‘ri sheng yue heng’ (日升月恒 — ‘as the sun rises, as the moon waxes’) — a kind of beautiful hope for the future.
Zhang Xiaojun: You’ve just described the overall line of thinking in the architecture design. For each chip generation — from the 910 A, B, C through to the 950 — what were the architectural changes and evolution in each generation?
Liao Heng: With the architectural changes, I think there were first some fairly clear and important realizations. The first: as I said, during those four to five years when we were stalled, our peers working on algorithms kept pushing forward — large models exploded, and they achieved rich results. Those results gave us a great deal of grounding for our designs.
The second very important event: some models — Llama being the first — chose the open-source route. As a result, roughly from its first generation, Llama became something everyone could analyze — it was like cutting open the body of a large model, letting people truly see what a large model looks like inside, down to every microscopic detail. That gave us excellent grounding for design. Later, of course, this series of Chinese models — whether DeepSeek or Qwen (千问) — kept open-sourcing, and in capability they were also close to the first tier, serving as that kind of reference point.
Beyond that, there was a lot in relatively non-language domains — for example China’s video generation and multimodal work, including Alibaba’s Wan series, and other internet companies, who also open-sourced a great deal of valuable work in those areas. Those works, plus our own closed business loops in autonomous-driving systems and in phones themselves, gave us a lot of real performance feedback, which also kept improving the models themselves. Simply put, by the time of the 950, the reference and evaluation basis for our designs was much richer, much more plentiful.
Second: after last year’s Spring Festival, when DeepSeek’s R1 came out, it directly triggered an explosion in model inference that we had never anticipated, and its scale was already quite considerable. So we gained a much more realistic understanding of the capabilities inference requires. Because inference doesn’t just have to be functional — it also has to be fast, right, the latency has to be low; and with low latency, the number of tokens each card produces — that is, the throughput — also has to be high. Balancing those several factors to the extreme gave us a sense of reality. Because before large-scale inference systems were deployed, we simply didn’t know these requirements existed.
Very quickly, once we started serving inference customers and genuinely deploying inference systems, we ran into a huge variety of engineering problems. One of the most visible: once you enter inference-system deployment — within maybe one or two months — you discover that although we used to think compute mattered most, in an inference system, especially in the decode part, compute actually doesn’t need to be that high; instead, memory bandwidth is especially important. As for the ratio between the two — if you take compute divided by memory bandwidth — you find, like the analogy I used earlier about the coordinated proportions of the human body at various scales, that the ratio needed for inference decode is markedly different from the ratio for training, for prefill.
And of course this process also, fortunately, confirmed one thing: what we call the supernode (超节点) — combining a very large number of chips into one tightly coupled, very closely cooperating whole — that capability is an important reason our 910C, and the 950, have been able to survive. Even though beforehand we had only a hazy vision of combining more things together, it was validated in actual business deployment. And of course we also discovered the places where we hadn’t done well.
Zhang Xiaojun: Which were?
Liao Heng: Namely: when you do communication — say, moving one megabyte of data at a time versus moving 7 kilobytes — the difference in granularity is an utterly fundamental issue. And the granularity of what we built in the past was all on the large side. What do I mean by granularity? If you fill a bucket or a wooden box with a pile of cobblestones, there are lots of gaps in between, because the stones’ granularity is large. If you fill the same box with sand, you find it packs much fuller, because each grain of sand is small. If you pour in a bucket of water — as long as the box doesn’t leak — it fills even fuller: that’s what I mean, because water is individual molecules, right, finer than any grain of sand.
So, for example, we originally thought this tightly coupled network would be used to move big blocks of data; later I found it was used to move very small blocks — we thought one megabyte, and it turned out to be 7 kilobytes. That’s a big difference, right? If the design isn’t good, you may be very efficient moving one megabyte of data but inefficient moving 7 kilobytes. So all these places need continuous polishing in practice.
So I can say why the 950 is significantly better than the 910C: first, that one is simply too old; second, that one was essentially designed by blind guessing, whereas by the time we designed the 950 there were a great many real-world things against which to test whether our guesses were right — and wherever they were wrong, we could quickly polish, optimize, and adjust. Moreover, the products have differentiated: as I said, the chips needed for inference and the chips needed for training require different ratios, and so may need different types of memory.
Of course, the last half year has brought an even deeper lesson. The lesson is — well, in fact we realized seven years ago that memory is especially important: rather than saying we sell an AI compute chip, you could better say we sell high-bandwidth memory. That realization directly drove us, seven years ago, to set up a dedicated department to build design capability for high-bandwidth memory. Yet even with seven years of preparation, when the storm arrived we found our preparation still wasn’t enough. In this current memory wave, you’ve seen prices rise more than tenfold, delivering an enormous shock to the entire electronics industry. What I want to say is: even with that much advance understanding and preparation, on the day the storm came — the day the tsunami came — there were still many regrets: OK, why didn’t we do more back then, put in more painstaking work, or prepare better in some way. Before people truly face the storm, they inevitably harbor some wishful thinking.
Zhang Xiaojun: Between chips needed for training and chips needed for inference, which do you think will be more numerous in the long run? What will the ratio look like?
Liao Heng: I think it should already be true today — this has already happened: inference will certainly be more. If inference isn’t the larger share, it means your economics are bad. Because inference is a revenue-producing process, while training is an investment in building capability: one is pure expense — a cost center — and the other is a profit center. And the profit center must be bigger than the cost center.
Zhang Xiaojun: In which chip generation will your thinking about these inference forms be reflected?
Liao Heng: I think it is strongly reflected in this very generation. For example, we split it into two tiers: one tier with very high bandwidth, and one tier with relatively lower bandwidth but better economics. For training, say, you might be able to use the more economical one.
Zhang Xiaojun: At the 2021 STW — the System Technology Workshop — you proposed that computing architecture had a new trend: moving from the old heterogeneous computing architecture, with CPU, GPU and NPU separated, toward a new type of homogeneous computing architecture. Could you explain what heterogeneous and homogeneous mean, and what distinguishes this new homogeneity from traditional homogeneity?
Liao Heng: To some degree I’d say that was something of a gimmick. A gimmick meant to push back against certain voices of doubt.
Zhang Xiaojun: Internal doubts, or external ones?
Liao Heng: There were doubting voices both inside and outside. First — hold on, this topic goes fairly deep. The first thing I want to say is that the most common argument in the industry is the SIMD-versus-SIMT fight. Some people say the GPU is SIMT — Single Instruction, Multiple Threads — that’s the GPU architecture; while Ascend currently — in the 910B and 910C — is a SIMD architecture. What does that mean? When you process an operation, do you express the computation as single elements, or do you compute on a big block of data at once? Put another way, it’s like the difference between a truck and a private car, right? If you drive yourself to work, each person takes one car; if you ride a bus, each bus carries a whole load of people — maybe dozens of people sharing one vehicle.
On the surface, this SIMD thing — what I want to say is, why is it a gimmick? In reality, today’s SIMT processors are internally SIMD as well; and today’s SIMD processors themselves also have multiple cores running almost identical code — that SPMD kind of pattern. But some people love to split hairs — or rather, people especially keen to exploit the CUDA ecosystem will feel SIMT is wonderful, because a huge amount of original research work has been developed on the CUDA ecosystem. So if your processor looks exactly like a GPU, you can download the code straight from GitHub and it runs without you doing anything — the ecosystem’s first-mover advantage.
But as I said, SIMT itself has SIMD elements inside it. For example, you’re processing an 8-bit floating-point number, but the compute unit may be 32-bit, so you have to pack four floats together to do the computation — that’s a little bus too. Nobody is fundamentalist about computing only one number at a time; if you computed only one number at a time, you’d be wasting the compute unit’s capacity. So even in code written for SIMT, people pack four 8-bit numbers together for computation. That’s an analogy.
So it comes down to nothing more than whether the bus is big or small. Undeniably, a small bus means small granularity, which helps with filling — as in my earlier example: filling a box with cobblestones versus fine sand, the sand packs fuller, right? So all we are really doing is judging whether a data block’s size should be large or small. The difference is 128 bytes versus 512 bytes: in effect, the GPU’s natural width is 128B, while our Ascend — especially the early products — had a width of 512B, which is a bit too wide. We have to admit that, to be honest and factual: our cobblestones were a bit too big. Although in the great majority of networks — especially Transformer-type networks, which are themselves very large-block, so a slightly bigger cobblestone doesn’t actually matter much — in certain networks — recommendation networks, or some early vision networks, especially on-device ones — smaller granularity may be needed.
In the 950 generation, what we call the new homogeneity means the processor supports both SIMD and SIMT, so the problem is greatly eased. Rather than carrying on this religious-style argument, better to use both — or to use different modes of computation in different situations. Because apart from the most essential point — the granularity question I just described — by today perhaps 90% of this issue can be set aside, since the technology as a whole has evolved to this stage; arguing about it further is no longer a core, critical question.
What matters instead is what I said just now: how much memory bandwidth to provision; how the interconnect should be built into a supernode; how to achieve extreme fusion — mega kernels — in kernel computation; how to reduce the synchronization overhead among multiple cores, even multiple chips. Those are the questions of the present — we consider the earlier debates to be already-solved problems, but these are the more important problems of today.
Zhang Xiaojun: These more important present-day problems — do you have answers now, or some interim thinking?
Liao Heng: Of course we have our own answers — otherwise the product couldn’t be carried forward. In our pipeline we have the 960, 970, and 980, all in the development pipeline, and we keep producing more answers of our own.
But let me try to tell you a more serious fact. In the past, AI algorithm developers were all used to PyTorch: to describe an algorithm, I use a pile of Torch operators — each operator is a function, a basic computational function — and stitching these operators together represents the whole model’s computation. But the more important present-day problem I mean is this: suppose you want to develop a modern model that way and also achieve, say, one-millisecond inference latency. One millisecond means: when conversing with the model — whether it’s a human or an agent talking to the model — submitting a prompt and getting output, one millisecond implies outputting 1,000 tokens per second. Today our typical chatbots — the great majority — output 20 or 30 tokens per second. Going from 20 to 1,000 is a 50-fold time compression.
Fifty-fold time compression — or, from a latency of maybe 20 milliseconds down to 1 millisecond, call it 20-fold — right, twenty or thirty tokens means forty or fifty milliseconds down to one millisecond, so a compression of several tens of times. At that point, the more important problem is: in almost every case where you are chasing performance and latency, you can no longer make calls operator by operator — because every call carries launch overhead and latency, plus the data-transfer overhead of the launch. So the more important question becomes: when these processors write the entire model as one enormous operator, what should that style of writing look like? That style is what we call a mega kernel, or kernel fusion: fusing tens or hundreds of operations into one larger operation, to reduce the launch and data-transfer overhead.
This change has been happening rapidly over the past year or two. The most famous work is Flash Attention; and the more important work after Flash Attention is the series of operators DeepSeek has open-sourced, and so on — every one of them an outstanding exemplar of fusion. This fusion process in turn delivers a big shock to processor design. In the past, when people designed a core, they only cared about one thing: my computation runs fast, and that’s enough. Now the question is: if I want to express a very complex computational process, the programming has to be reasonably easy to write, and I have to be able to modify the program on your processor well and quickly, so that it achieves maximum performance. This has given rise to some modern programming approaches. People aren’t quite — especially when the algorithms are changing very fast — people can no longer bear expressing such complex computations through low-level programming.
So naturally, new programming languages and compiler technologies emerge. These include, as mentioned earlier, Triton, which is a piece of OpenAI’s work; then TileLang, built by Professor Yang Zhi of Peking University and his student Wang Lei, which is currently also DeepSeek’s main development approach; and PyPTO, which we built ourselves. All of these try to express extremely complex computation in a higher-level way — to the point that an entire model is written as a single operator. And that has a fairly far-reaching influence, from the software side, on compilation and on processor design.
Because if you look at the classic Hennessy and Patterson textbook, Computer Architecture — the first time I saw this sentence, I was genuinely, deeply struck. It says computer architecture is ‘the interface between software and hardware’ — it is that interface, right —
Zhang Xiaojun: The communication interface between the two.
Liao Heng: The communication interface between the two. Did most of the people who used to work on processors, on the hardware core side, really have a particularly deep appreciation of this? Because they only saw their own side, the core: I need to design a compute unit — how do I make it efficient, how do I save die area. But once you raise your understanding to the level of this interface, you can see that changes in programming languages affect this interface — it is two sides of one thing; the interface necessarily has its yin side and its yang side — and where the dividing line falls is determined not only by the core, but also by the way you program and the way you compile.
So for the design of today’s AI cores — of Ascend processors — I think we, or rather our peers, everyone who designs processors, are facing the fact that this interface is being reopened. Or to put it another way, the protective wall of the CUDA ecosystem is rapidly disappearing, because it is no longer the primary interface. If all of DeepSeek’s work is built on top of TileLang, then TileLang becomes a critically important interface.
Zhang Xiaojun: So that’s also the point of what you’re building now?
Liao Heng: Right — we ourselves of course had this understanding years ago. Despite facing a great deal of skepticism, we have also been doing our own open-source compiler work, the PyPTO line of work. This work is purely first-principles based; there isn’t much that insists on… that is to say, the methods and design philosophy we use should be universally applicable, not applicable only to Ascend. In other words, although it was we who open-sourced this work, I believe all AI processors can benefit from the same compiler stack — that is, what’s called cross-platform.
Zhang Xiaojun: CANN — what are your expectations for it?
Liao Heng: CANN is our software stack — essentially our entire suite built on top of Ascend: the runtime environment, compilers, libraries, plus the models we have already ported.
Zhang Xiaojun: Do you expect it to be the next CUDA-type product?
Liao Heng: It is already comprehensive — on one hand, it is comprehensively protecting the legacy. By ‘legacy’ I mean: whatever capabilities CUDA has, our CANN must have too — even one-to-one at every single layer. For example, you can run PyTorch, I can also run PyTorch; you can run vLLM, I must run vLLM too. But we genuinely still have a certain gap — that’s the legacy part. Then there’s the part that is advancing by leaps and bounds.
First, legacy serves the masses. What do I mean by the masses? If you’re a university student just starting to study AI — the deep neural networks course — you’ll most likely use PyTorch, use Torch operators. If the professor teaching the course uses Ascend hardware as the environment for the coursework, then you’ll need CANN. All of that belongs to legacy; it’s the long tail, and it may persist for a very long time.
But the part serving the few is this: when I want to deploy a GLM 5.2 and I want the highest economic return — the most tokens produced per card per second — and I also want the lowest latency, that’s where the ‘advanced’ part comes in, right? First we need to use more advanced methods, in the shortest possible time, so that people can quickly achieve this high-throughput, low-latency tuning. And this tuning is tuning taken to the extreme, because every additional 10% of throughput is an additional 10% of revenue, right? So it’s real money. And that advanced side may not involve ten thousand PhD students using the thing — maybe only a hundred people working day and night to crack the problem, raising the throughput of the production system as early as possible, because it is a production system. I think we are making continuous efforts on both fronts.
The reason Ascend can sell at all is, first of all, that both sides need to show progress. If the performance of the production system doesn’t go up, nobody will buy it, right? Because if you buy it, you lose money — you spend the same amount of CAPEX and produce fewer tokens; or you train a model and it won’t converge, or it takes longer — converging five times slower than others — nobody can accept that, right? And the ecosystem part I just mentioned serves the broader community. For this community, we have in the past gone fully open source and worked with countless university professors, hoping they will build this ecosystem together with us. The CANN community has now become the hottest open-source community in China, because so many people follow AI and there is a great deal of enthusiasm around it. So, riding this ‘many hands make the flames rise higher’ trend, we have made great strides over the past year.
Zhang Xiaojun: Which do you think is harder — this extreme physical pursuit at the nanometer scale, or building an ecosystem that surpasses CUDA? Which is harder?
Liao Heng: I would say surpassing CUDA is harder — even though the extreme pursuit at the nanometer scale, or making breakthroughs in physics, is also hard. Why? One of them mainly rests on your own effort. The other requires changing the shared habits of a whole group — of countless people. For example, if today you invented an instant-messaging app and tried to persuade everyone to stop using WeChat and use your new app instead — that’s hard, extremely hard, right? Because it’s a group habit, and that group habit needs to keep receiving positive reasons before it will make such a shift. Whereas developing the next, more advanced chip, as I said — as long as we ourselves put in the effort, maybe 70% of the factors are in our own hands, and perhaps 30% are objective physical constraints, and we just have to find every possible way to break through those physical constraints. But these two problems are, in essence, solved by colleagues from different disciplines.
Zhang Xiaojun: Solved by them.
Liao Heng: Right.
Zhang Xiaojun: Was open-sourcing CANN a difficult decision for you?
Liao Heng: It was not a difficult decision. Why? Well, I have my own personal view — of course everyone has their own perspective. First, communication between humans: the human brain is highly developed, but the connection capacity between people is very poor. For example, sitting here doing this interview with you, my speaking rate is maybe just three to five characters per second — and that already counts as fast, right? So the communication bandwidth I give you may be only 1 kbps. But if you were an Ascend chip and I were an Ascend chip, our communication bandwidth would be 7.2T — 7.2 TB per second. TB per second is a terrifying number — who even knows what power of two that is, right? So the orders of magnitude between the two differ enormously.
So why do I use this metaphor to describe the importance of open source? Because when you deliver a very complex technical system, what’s actually needed is for the other side — the people using the thing — to quickly understand: what are this system’s interfaces, how is it designed, how should it be used? So this communication bandwidth — when we go and restrict such a communication bandwidth, you realize that this communication was already an enormous bottleneck to begin with; and when on top of that you control when a given document may be sent to whom, and even viewing it requires signing an NDA — that inherently adds unnecessary security gates onto an already very constrained communication channel, right? So-called open source is really openness taken to the extreme. Meaning: once I have shown you all the code, I actually no longer need to communicate with you, because you can go look for yourself at the countless files in the countless directories, right — no more need to waste breath on it. So it solves the biggest bottleneck in this kind of exchange — sharing ideas and getting aligned.
Liao Heng: So we feel that from the moment we open-sourced, both on our side — the developers — and on the customers’ side, there was an immediate sense of a weight being lifted. A problem that used to be difficult — every time you had to make phone calls, file issue tickets — now that man-made bottleneck has disappeared, right, and the speed of solving problems is much faster. And that’s not even counting the many more people making their own contributions on top of it. There’s also the fact that AI-assisted programming is now becoming an important — even a dominant — way of working for us. After open-sourcing, those models and agents can much more quickly get an ample corpus, giving them training data for programming on CANN. So those models are also playing a very good role in CANN development.
Zhang Xiaojun: I feel the early Ascend story had something of a backed-into-a-corner, rising-up-in-resistance flavor to it. Coming to today, do you think Ascend and Nvidia are becoming more and more alike, or less and less alike? Have you taken the same path or a different one?
Liao Heng: I would say that if there were more similarities early on, part of it is that they have become more and more like us, right — the Cube I mentioned earlier: their tensor core has come to look more and more like our Cube, getting bigger —
Zhang Xiaojun: Bigger.
Liao Heng: — to achieve higher compute efficiency. But apart from such details, at the macro level things were fairly similar back then, because everyone was centered on the single chip, and deployment approaches and so on were all fairly alike. The bigger dissimilarity, though — I think we are now becoming less and less alike. The most fundamental reason lies in what I described earlier: the difference of small versus large, the difference in granularity, the difference of using the small to fight the large, the difference of the supernode — including various system-level differences — less and less alike. At the micro level, we may still have plenty to learn from them — for example, as I said, realizing our granularity had become too coarse and should be made smaller, roughly their size, so that adaptation efficiency across more algorithms is higher. But at the system level, we are less and less alike.
Let me offer one small angle to try to illustrate this. At least — this is our bold speculation, or judging from the roadmap they announced at GTC — I think Nvidia has been pursuing extreme increases in rack-level density. The increases are astonishing: almost every generation is a doubling — twice the compute stuffed into the same rack. Personally, I think there has to be a limit; you cannot push it to the extreme, and you should not increase rack density without bound. Why? There are several dimensions to this. First, I think the specs of a single chip should not be pushed to the extreme — or simply cannot be pushed further. The single chip — OK, the single package. Why shouldn’t a single rack push its specs to the extreme? It does need to improve, but it should not exceed the limits of physics. Because once you exceed the physical limits, you run into catastrophic consequences.
Let me start with the single chip. Everyone has seen that one of the biggest predicaments the AI industry has created is the HBM shortage — because demand has exploded, everyone is using bigger, more numerous, higher-bandwidth HBM. Let me try to explain, with a fairly vivid example, why this kind of 2.5D scaling has a ceiling. Take the compute die — there are actually multiple dies these days, but let’s simplify the problem — suppose I treat the compute die as a square with side length N; then its area is N squared, right? A square with side N — isn’t its area N squared? Then its compute power is N squared — because each compute unit requires some number of logic gates, and if you have N-squared gates, compute power is proportional to N squared. But when you have N-squared compute power, the memory bandwidth and interconnect bandwidth you consume scale in the same proportion.
But if I place one huge compute die in the middle and keep making N bigger, while my HBM, all my IO, and all my power delivery are arranged around its perimeter — then doesn’t every single bit of memory access have to cross the edges of this square’s four sides? The perimeter of the four sides is 4N. So you find that as N grows bigger and bigger, your compute power, your bandwidth, your interconnect, your power consumption all scale as N squared — but your four sides only give you 4N. Between that quadratic curve and a straight line, won’t there be an ever-widening gap? That opening — the contradiction between the two — will explode more and more, right? So this shows that the current path of relying on ever-bigger compute dies plus ever-more HBM has a ceiling. Of course, before we reach that physical boundary, we will keep making it bigger; but once you cross that boundary, you will see catastrophic consequences — an avalanche.
Zhang Xiaojun: How big might that boundary be?
Liao Heng: The boundary might be four reticles, each reticle being 800 square millimeters; maybe six reticles, maybe eight. I trust our whole upstream — and this gets into the lower levels, the manufacturing chain I described earlier, the colleagues on the fifth, fourth, and third floors, the people in the basement — to see how far they can extend that capability boundary. But I already know with certainty that this physical gap between N squared and N is irreconcilable. So this kind of architecture can perhaps run for one or two more generations, maybe three, but most likely by the fourth generation it will no longer work.
Zhang Xiaojun: Because they are still building along this architecture?
Liao Heng: Currently, yes — the roadmap they have published so far is like that. So I don’t think a single chip can keep doubling every year. If you double every year, you make this gap ever more contradictory, ever bigger, right?
The second question is: why can’t a single rack be scaled up without limit — and why is it also unnecessary? Because if a rack’s power draw is 100 kilowatts today, 200 kilowatts next year, 400 kilowatts the year after, and the year after that — reaching a megawatt might take only four years — two to the fourth power, already many times over. Now why shouldn’t you push it that high? Because as you stuff in more and more chips, they have countless interconnects — everything — there are chip-to-chip links, and they need bandwidth. The more you stuff in — it’s not just the chips’ compute that doubles: their interconnect doubles too, their power draw doubles, their heat dissipation doubles. Let’s look at it in mundane, practical terms: for a one-megawatt rack, if it used natural ambient-air cooling, we can do a back-of-the-envelope calculation of how much cooling-tower space is needed. Say it takes 500 square meters to dissipate one megawatt — then you find that a gigawatt data center might contain only 1,000 racks, yet it might need a full square kilometer of land just to house the cooling towers.
Between those two — a rack’s footprint may be under about two square meters, but if it needs 500 square meters of space for cooling, that’s a 250-fold spatial amplification. When that ratio becomes seriously out of proportion, you find that the water pipe may have to run a kilometer to reach the cooling tower, because the space the cooling towers occupy is vastly larger than the space your machines occupy, right? Everything I just described actually goes far beyond the elements of processor architecture design itself — but it belongs to what I called another layer of the problem, right? In terms of the 18-layer pagoda: if you build an AI gigawatt data center, someone always has to construct the buildings, lay the water pipes, design the cooling system. The point is that this ratio must not become seriously distorted. We can only say — it’s really not our place to comment on how others do it — we can only say that in our own system design, perhaps as early as three years ago, we had already realized: I do not want to stuff too many chips into the same rack.
That then means the interconnect distance between my chips — if others stuff 100 chips into one rack while I spread them across 4 racks — the physical distance between those 100 chips becomes longer. Of course, longer physical distance brings latency and all sorts of overheads. But if I recognize that this is a fundamental problem — that one day I will have to spread things out — then I need a wiring approach whose connection cost is very low, and whose transmission distance won’t get to the point where the signal can no longer get through. So we were in fact proactive: when we designed the first-generation UB — the UnifiedBus (灵衢) I mentioned earlier — we designed it so that it can run both over electrical wiring and over optical cable. And you should understand that the thickness of an optical cable and of an electrical cable differ greatly — the diameters may differ by a factor of ten. When you’re connecting a single wire, there isn’t much difference; but when a rack has to connect five thousand wires, you see the physical reality — the size of five thousand electrical cables bundled together may already be as thick as an elephant’s leg.
I think Nvidia’s approach is rather brute-force. On one hand they keep raising speeds — pushing up the SerDes rate, the signaling rate on every single wire. Everyone is raising speeds; everyone wants to carry more signal over fewer wires. But on the other hand, they have used the most advanced board materials in the world, which has directly caused a shortage in high-end PCB capacity too; and they have also used some very advanced physical-layer technologies, right? All of this is fantastic — these are areas where they are extremely good. But I, as a system designer, do not want to take on the world’s very hardest problems in every single spot. Why? Because with every generation, I have to deliver a new, highly reliable product to customers, on time and at quality. If I have to break through 20 physical limits simultaneously, then if there is even one I fail to break through, my product is dead. I would rather put my energy into breaking through maybe five especially hard — and high-value — problems, and leave some of the other hard problems for others to solve. That way, my probability of failure — of the product dying — drops substantially, right? Because even if each one has a 99% probability of success, 99% to the 20th power is a number approaching zero; whereas 99% to the 5th power may still be tolerable. So this is a design philosophy — a philosophical question: play to your strengths and avoid your weaknesses, bring what you are good at into full play, and do not try to contend with others on every front.
Zhang Xiaojun: So on the question of whether the single chip keeps getting bigger and bigger — is your answer that…
Liao Heng: Of course we are making it bigger and bigger, but we do it within our own comfort zone.
Zhang Xiaojun: Not unboundedly big.
Liao Heng: Not unboundedly big. We will not go and take on the kind of design that is simply asking to get itself killed — we won’t do that.
Zhang Xiaojun: And single-rack density — should it be pushed to a megawatt?
Liao Heng: We don’t need it — the answer is no. We know with certainty that we don’t need it. Because our space is dirt cheap — each square meter of a machine hall may cost only 2,000 RMB. Compared with a 10-million-RMB rack, that 2,000 RMB is completely negligible. So we have space in abundance. Through optical interconnect, we let these chips sit comfortably spread out across a larger space, and thereby obtain data centers of 100,000 or even 500,000 cards — we can build all of that. A single machine hall can probably reach 100,000 — 100,000 chips is already no problem at all today.
Zhang Xiaojun: In the large-model industry, people still say we’re roughly a few months behind. In the chip industry, what do you think our real generational gap is?
Liao Heng: I think on this question — people have this worry, or…
The publication of the Tau Scaling Law (Tau定律) I mentioned earlier was, to some degree, a way of letting more people know: shrinking (微缩) is not the only path. When you can no longer shrink, you can stack — and stacking can be done at the chip level or at the system level. I believe that in the second half of this year, when people see the newest phones —
It should be released in the autumn. Because there are plenty of outfits like SemiAnalysis and TechInsights that will certainly do reverse engineering, you will see a lot of analysis reports come out, and you will be able to see whether the so-called Tau Scaling Law, stacking, and logic folding can actually solve — or narrow — the generational gap. My answer is: to a large extent, yes. And people may have consistently overlooked the most critical point of all: when people buy a GPU or NPU, they are actually buying memory. Because calculated at today’s costs, 80% of a GPU or NPU is the cost of the HBM; only around 10% of the money is spent on the advanced logic die. So the main competitiveness of an AI chip arguably comes from how much bandwidth it has, not from how large its compute specs are — or rather, when the compute specs grow, the bandwidth must scale up in proportion. After all, HBM cannot be sold on its own; it has to live parasitically inside some GPU or NPU, so its monetization depends on the final integrator. The reason I bring this topic up again is that the biggest core competitiveness does not necessarily come from how advanced that logic die is — it comes more from how advanced the aggregate bandwidth and the memory matched to it are, right?
Zhang Xiaojun: What you have just given are a lot of conclusions. In the actual process of feeling your way forward, when did you find this path?
Liao Heng: I would have to say our capability in this respect is still somewhat lacking. The first thing is instinct. If you are lost deep in the mountains and forests, a good hunter has a kind of innate intuition about which direction might lead to water to drink, which direction leads back home — maybe he looks at the stars in the sky to judge his bearings, then picks a direction and goes. So-called instinct means making a choice without sufficient evidence — intuition, plus some judgments of your own. And those judgments are not necessarily rigorous conclusions from reasoning that could be mathematically proven, right?
The second thing, of course, is that you need to keep seeking verification in a given direction, and run some experiments yourself. And even more — this is where I say we are still lacking — we need to find the people building models, or working at the scientific frontier on all kinds of quantitative work, and form an effective dialogue with them: our instinct is such-and-such, your work is in this field, let us see whether we can align. So this is a very important part of the four-year lead time I mentioned earlier: first, you yourself need enough credibility that others are willing to engage with you and form an effective exchange; and then, at that level, people need a kind of exchange of ideas, and continuous alignment.
Zhang Xiaojun: A lot of capability comes from those model pioneers. Can you talk about some stories from the past two years of mutual support, adaptation and optimization with some Chinese large models — for example, DeepSeek?
Liao Heng: The specifics I am not in a position to discuss. But what I want to say is that in the so-called AI ecosystem there are perhaps three different levels, with the difficulty and the requirements rising step by step. The first level is: a model arrives, and I can run inference on it. That is relatively simple. When a model arrives, its service life usually lasts — maybe half a year, or perhaps a year. Once you have adapted it and tuned its performance well, people will copy and deploy it over and over, right? That is inference: one-off work that can be replicated.
The second level: taking a model that a research team has already finalized into first-version production training — the workload is perhaps three times that of adapting an inference system. All the operators already used in inference are needed in training too, but there are also backward operators, and backward operators are usually a bit more complex. On the other hand, the demands for extreme performance efficiency are relatively lower than for a production inference system. Either way, large-scale training is also one-off — or rather, you do it once and it runs for three months, then you do the next version. This still counts as a difficulty that can be solved, and crossed, with manpower.
The third stage is research. An advanced research team may have dozens of excellent algorithm people, each running different experiments every day, constantly modifying their code. We consider the inference breakthrough basically solved, and the training breakthrough entirely within reach — it will not take very long. But breaking through at the research level can only be accomplished with a more advanced compiler stack. Researchers have now switched to TileLang or other higher-level languages, because they are easy to modify. No matter whose processor you use, everyone actually faces the same problem: change the algorithm a little and you have to rewrite the CUDA operators — nobody can stand that. So we are also working very hard on this with compiler stacks — no matter which company leads them — we are striving to get this right. Once that is done, we will be in a position to expect that within the next year or so we may partially reach that research level I just described. But it genuinely takes time.
[3:00:00]
Zhang Xiaojun: In your view, as of today, is there original innovation in your work? Where does it show?
Liao Heng: Perhaps I can talk about it on a few levels. One is the chip itself. I personally am a staunch member of the ‘don’t copy the homework’ camp. Not copying, of course, creates enormous difficulty, because not copying means you cannot borrow someone else’s ecosystem — you have to pay that price. But I have a deeply rooted conviction: we are a technology company, and no technology company in the world has ever achieved great success by copying homework. Because technology itself is innovation — your primary means of value creation is building a differentiated advantage. Even if you have a hundred flaws, you must have one strength. That is my understanding after working in this industry for so long. I am especially afraid of making products that are the same as other people’s. Because we are not the Haining Leather City or the Wenzhou small-commodities market — where every stall’s Christmas trees look exactly the same, and then it is just a contest of who has the lower cost. That model may work in low-end industries, but in an industry that leads the world it is absolutely unacceptable. You must have your own strength, because only that strength converts into the value of the product.
Zhang Xiaojun: When was ‘no copying homework’ decided? From day one?
Liao Heng: From day one. Or rather, in every product I have ever worked on, we have never copied anyone else’s homework. Put another way, the moment I hear we are supposed to be the same as someone else, I feel that definitely will not do.
Zhang Xiaojun: What is the price of not copying the homework?
Liao Heng: The price of not copying is, for example, that your ecosystem gap is enormous at the start. Everyone is used to Windows, and you suddenly want to make a non-Windows PC — you have to change those people’s habits. Or, like now, everyone uses iOS or Android, and you suddenly want to build HarmonyOS (鸿蒙) — you have to put in enormous effort to establish a unique advantage, or give people a reason to switch. That is the price. As for these problems, I think we are now crossing the so-called painful plateau, and may be entering a better phase.
Zhang Xiaojun: For you, would copying the homework not be easier than not copying?
Liao Heng: Of course it would.
Zhang Xiaojun: Then why insist on not copying — from the standpoint of your position?
Liao Heng: Because our mission is not to win some praise — or take a scolding — in an exchange with some customer this year. Our mission is to build a sustainably developing product system. And within that product system — as I mentioned from that Hennessy and Patterson book, the interface between software and hardware is two sides of one coin, right? Copying the homework means: I want to make my hardware sufficiently similar to someone else’s, and then for software I entirely use the ready-made software others have built, so I do not need to put in much effort building my own ecosystem. But I am deeply convinced this has several fatal problems. The first problem: you can never — suppose today I want to transform myself to become exactly like you — I can never look completely identical to you, and there will inevitably be consequences from not being exactly the same.
Second: such a technical system is completely incapable of evolving. Because if you copy step by step, you must copy at every step; you have entirely lost the possibility of independent design innovation. The other side’s software is already there, so you have to cut the foot to fit the shoe — whittle your own shoe down until it fits into, what, Snow White’s crystal slipper (as spoken — the glass slipper is of course Cinderella’s), right? Once you have whittled once, next time you have to whittle again. So I think this is the follower’s inevitable backwardness: the very process of copying along has already determined that he will be behind. And we do not want to be behind — we want independence and self-reliance. Moreover, our problem, our constraint — say even if I copied to be the same as him, his interconnect capability is different from mine. What do I do then? There will inevitably be a communications part I cannot copy. And once the communications part cannot be copied, the operators that fuse communication and compute cannot be used. So with this kind of problem, it is not enough to copy part of it — you would have to copy all of it.
Zhang Xiaojun: We were just discussing innovation — you covered chip architecture; what about other aspects?
Liao Heng: As I said, the important part is the system level. As mentioned, we have completely taken two different roads, with different design philosophies: they go dense, I go sparse — right, I go loose. I want everyone to have their own bedroom and sleep comfortably; he wants twenty people sleeping on one kang (heated brick bed).
Zhang Xiaojun: For the next few chips that have not yet been officially released — including the 950DT, and the future 960, 970, 990 — what should we look forward to?
Liao Heng: I think the foreseeable next three generations are all about the same — roughly on par with the competitor — at a cadence of doubling every year. Beyond that, there is the system level. What I find most worth looking forward to is this: we have been pondering a question — do you build one highly integrated building block, or do you build smaller building blocks and, by wiring them together, let different users in different scenarios obtain more flexible combinations? That is a rather interesting system-level question.
Look at the equipment people used in data centers in the past: earliest were servers, each a 2U server; later, AI servers were maybe 6U or 8U — taller, the box is bigger; and then the Pod appeared. A Pod is basically the typical all-in-one: the whole rack frame is one device, with different switch boards, compute boards and a backplane on the front and back. These two forms obviously each have their strengths. First, the all-in-one device is highly integrated and can perhaps achieve extreme density. For instance, Nvidia’s NVL72 is an all-in-one big machine, right — and on the later roadmap there are bigger and bigger machines of this kind. That is an important form factor, and we have similar forms too. But I think while this form’s advantage is high density and the ability to pre-integrate, the challenge it brings is: if you want to make any adjustment to the proportions of what goes inside, you cannot. Because the machine has effectively hard-fixed all the parts that can be plugged in, the ratios between them, and their connection relationships.
And what we see is precisely that, because of the business itself, models went from 600-odd billion parameters to 1.6T in a single year, and next year it may be 5T to 10T. The resources they need — the KV cache used to sit in the CPU’s DRAM, and now everyone has migrated it to SSDs; and model inference speed, or latency, has to drop from 20 milliseconds to 1 millisecond. Whether in size, speed or latency requirements, these are drastic changes of half an order of magnitude to a full order of magnitude. Such changes may well be hard to predict at the time you design the Pod — two years later, when the machine reaches the market, the original proportions have already shifted a great deal. What then? So going forward we will quite actively consider the need for a more reconfigurable kind of building block, where the blocks can be recombined between themselves in the simplest possible way. This actually ties in with the relatively loose design philosophy I described earlier, and with using all-optical interconnect — these ideas are all connected, or mutually reinforcing. Put another way, once I have all-optical interconnect, I can build smaller boxes, and these boxes can be freely combined — just by connecting a few optical fibers — into the optimal configuration. So that is a system-level innovation, or something to look forward to. And our supernode domain will also be made very large, because we have found that if you have a 20T model requiring ultra-low latency, the communication domain becomes quite large — hundreds of nodes, even a thousand-plus nodes.
Zhang Xiaojun: In the chip field, would you say China has walked out a Chinese path — in your view, as of now?
Liao Heng: I do not think my saying so actually counts for much. I think you can go look at these companies’ financial reports. Their financials to a large degree represent an aggregation, because the great majority of the companies are fabless — everyone goes to certain fabs for manufacturing — so it is roughly equivalent to a sum of the whole industry. The financial reports are very telling, because they represent, first, a scale of economics, and second, whether they can be positively profitable. Because whether a Chinese path has been walked out requires not only technological breakthroughs but also an economically closed loop. You cannot subsidize forever, or bleed indefinitely — that is an unsustainable state, right? So I think if you look at those places, you will get a very good answer. And the financials are not just the logic fabs alone — you should also look at the memory fabs, the NAND fabs, and the packaging houses. My direct or indirect impression is that, especially on a five-year timeline, they are all in a state of order-of-magnitude surge.
The next question is whether that curve is sustainable, right — that it does not fall back down. I personally believe it will not. Why? Because first, in many fields we are actually not behind: advanced packaging is not behind; the NAND field, for instance, is actually not behind, right? The DRAM gap is not that large either, nor is the HBM gap — there is some gap. For logic, as long as the loop can close positively, that is already very good. And then of course there is the macro question: will China and the US ever return to that past so-called strategic-partner relationship of mutual trust and mutual willingness to depend on each other? I think it is quite clear the problems lie more on that side, right — the other side is not willing.
Zhang Xiaojun: Everyone says compute is scarce, compute is scarce — every model company is severely short of compute. So what is the true state of China’s compute today?
Liao Heng: I do not know what the actual situation is, but I can guess from some tangential data — and it is probably a very inaccurate guess. Reportedly, last year China deployed something over one gigawatt of data centers, while the US deployed roughly seven to eight gigawatts. So we are at roughly one-fifth — or somewhere between one-fifth and one-seventh or one-eighth — in relative terms. Of course, they have also deployed a lot in many other odd places owing to particular factors — Southeast Asia, for example.
Zhang Xiaojun: Can your supply volumes be ramped up?
Liao Heng: We are working hard on it — probably yes.
Zhang Xiaojun: Say by this time next year — how much more than this year?
Liao Heng: That is hard to say. All I can say is that this problem should still be solved rapidly.
Zhang Xiaojun: When did you realize you might have made it through the hardest days? Which year?
Liao Heng: I felt I had made it through the hardest days when I started receiving more and more criticism — especially internal criticism.
Zhang Xiaojun: Why?
Liao Heng: Because at the most difficult time, no one comes to criticize you.
Zhang Xiaojun: Oh.
Liao Heng: Last year, the year before — I would say by around last year or the year before, that point had already passed.
Zhang Xiaojun: So you were still being challenged internally?
Liao Heng: Why do I call this a way of drawing the line? Because if you are walking in a desert and have not drunk water for seven days, then any water you see, no matter how dirty, you will feel it is saving your life — you can drink it, and you will not be picky about how the water tastes. After the first bottle of water goes down, everyone’s body immediately recovers; then by maybe the second bottle, people start to feel the water tastes bad — and that is the moment of criticism I am talking about. So I think the most difficult time is when nobody criticizes you.
Zhang Xiaojun: Does being criticized feel unfair?
Liao Heng: I think, for me — perhaps my virtue is insufficient, or it is a flaw of temperament — there is still some pushback in me. But afterwards I come to think that this too is one of life’s inevitabilities, so there is nothing really to feel wronged about.
Zhang Xiaojun: You described those dozen-plus years at that overseas semiconductor company as living through a long period of decline and shrinkage — the sunset of a sunset industry. What about these last ten years? What have these ten years felt like?
Liao Heng: These ten years, for China — or for the environment I am in — even though we went through a nine-deaths-one-life calamity, have absolutely been an explosive period for capability.
Zhang Xiaojun: What has it felt like?
Liao Heng: The feeling is that, apart from process technology being a tiny bit behind — a little behind, specs a little behind — the vast majority of the chips America can make, China can now make itself, from design through manufacturing. And not just our own team — this includes the overall capability outside Huawei too. For example, you will see countless companies making autonomous-driving chips, or embodied-intelligence chips, or WiFi chips. You will find this capability has become fairly widespread — that is, the ability to design relatively complex SoCs with a certain barrier to entry is now blooming everywhere, which is to say this capability has become relatively less scarce. The second interesting point is that the age of the people in the field is a good twenty years younger than the Silicon Valley crowd.
Zhang Xiaojun: Twenty years younger? Than people in the Silicon Valley semiconductor industry?
Liao Heng: Yes. In Silicon Valley semiconductors I would count as at least not old — on the younger side; but in this industry in China, I am at an age that should probably soon be phased out. So this place represents a kind of hope.
Or put differently — young people have more explosive energy, and they have many more years ahead of them in this profession.
Zhang Xiaojun: So here is the question: so many people are now entering the chip industry — including the large-model companies. How do you view this domestic competition?
Liao Heng: In any case, I’m not particularly worried. I think having competition is quite healthy. I hope they are all very capable — but I am not afraid of them.
Zhang Xiaojun: During this ordeal, was there a so-called near-death moment? When was the most desperate moment?
Liao Heng: I didn’t experience one personally, so I don’t want to put myself into other people’s roles and recount those events on their behalf. For example, perhaps the closest thing to a near-death moment — one I was only indirectly involved in — was our Kirin (麒麟) chip, that is, the phone SoC. But fortunately there were some very resolute people who absolutely refused to give up, and they brought it back to life.
Talent and Compute
Zhang Xiaojun: Most of the companies I interview lean toward software. In a hardware-leaning industry like yours, is the design of the organization and its culture very different from those software-leaning companies? Do you have any distinctive culture?
Liao Heng: This topic is a bit deep. I can only say that, since I’m not someone responsible for talent management, or for managing a very large organization on the people side, whatever I say about talent from an organizational angle may not be very accurate. But I can, from a personal angle, describe how a talented person grows within a chip team — some points that are perhaps especially valuable and important, and how to grow. Maybe that’s more meaningful.
First: to be a chip engineer or architect, you absolutely must have actually done a chip. A genius PhD cannot do chips without that. You have to start from the module level — module design, subsystems — and go through this long process, typically at least 18 or 20 months, and you must go through it more than once — many times. Why? Because chip design is very different from writing software. When writing software, out of habit, after writing say ten minutes of code I will definitely run it: hit compile, then execute the code. With a chip, you only get to press the button once, after 18 months. So you have absolutely no room for trial and error — the moment you press that button, you have spent at least 200–300 million RMB.
Second: a chip involves a great many physical things. A fairly smart person is often strong in the digital world, in logical thinking — but you must go through the physical side: be off by one nanosecond, miss timing by one picosecond, and the chip immediately fails; it dies right in front of you. So there is this extremely rigorous, exacting engineering dimension, and school education cannot give you these things. You must have been through the process to deeply appreciate why you have to be so meticulous about this — because you cannot be off by a single picosecond, cannot route a single wire wrong, and a logic gate cannot contain any critical bug. Software, by contrast — you can just fix it a minute later, right? You find your bug when the run fails.
So I think this process is exactly why I am not especially worried about those seemingly star-studded teams — because perhaps their leaders or organizers have not realized this point. They put too much weight on how smart people are, how good their resumes look, or whether their education was under famous masters. It has nothing to do with famous masters — you must go through the experience, and that experience carries a cost. It is through a continual tape-out process that large numbers of people get forged. Someone unwilling to go through this process will never grow into the position of architect. That is the first point.
The second point: if we are doing chip design — say fabless design sits on the seventh floor of the stack — then do you have sixth-floor capability? What about the eighth and ninth floors? Can you reach a shared understanding with someone who works on algorithms? If you have no basic understanding, people won’t even bother engaging with you, right? If you cannot understand a word of what they are saying, the exchange simply cannot continue.
So I think, as an engineer, you need to open your field of view — your horn of curiosity — fairly wide: care not only about what is in your own hands, but also about what the people on the floor above are doing, and the floor below. And don’t just listen to others — you should Google things more; or nowadays, with large models, you can ask a lot of questions. For instance, I just asked one: how much cooling-tower area does one megawatt require? For a question like that, asking one model may not be enough — I will ask four in a row and compare whether their answers roughly agree in order of magnitude before I believe them. You see, these are not questions that someone focused solely on delivering their own assigned duties would ask — but it is a very meaningful question, isn’t it?
So what I want to say is: opening the horn wide means investing a lot of extra curiosity; and beyond curiosity, investing a lot else — reading a few more papers that seem to have nothing to do with you. Then, when you are able to hold a conversation with people working on the most frontier models, you will find you have entered another comfort zone: not only can you understand them, you can even predict what they will be thinking next, and you may even anticipate many problems they haven’t thought of — because you have a more low-level perspective, right? Then, given time, as your experience shuttling between different floors accumulates, you integrate more and more layers, and your ability to get a grip on a technical problem grows stronger.
Zhang Xiaojun: So when you value a candidate — setting the resume aside, someone who perhaps has not yet actually done a chip — what is it you look for in them?
Liao Heng: The first thing I look at is their values — or maybe using the word ‘values’ is putting it too strongly — basically, whether we can get along with one another. That said, maybe ‘values’ isn’t wrong either: whether we fit when discussing things, be it technical questions, expectations for the future, or expectations about the career. Because even a candidate whose other qualities all look extremely high — if they only spend half a year or a year, and leave before going through, without the patience to go through, those training cycles I just described, then for me it is a complete waste of time. So we increasingly look for a rough alignment of basic outlook — only then is it meaningful. Otherwise, we are just wasting each other’s lives.
Zhang Xiaojun: You don’t poach people with astronomical offers, do you?
Liao Heng: Astronomical offers to poach people — we don’t have that kind of capacity. We don’t have the capacity for sky-high prices, and it also wouldn’t really fit Huawei’s so-called ‘striver’ (奋斗者) culture.
Zhang Xiaojun: With a striver culture like Huawei’s, how does the organization create innovation? Or is it a militarized style of management?
Liao Heng: No, no — Huawei is not like that. I think the innovation lies entirely in things like this: an employee of mine who joined just a year ago can come to my office and argue an issue with me — and I very much welcome people like that. I can’t speak for the so-called organizational culture as a whole — I can only say that each of us sees only a very small range around ourselves, right? We work hard to help the people around us become self-driven — self-driving is more effective than being driven by others.
Zhang Xiaojun: There are also young people now facing the question of whether to build their careers overseas or return to China; a common concern of theirs is that resources at home are constrained.
Liao Heng: Which resources do you mean?
Zhang Xiaojun: Compute resources; the gap is large.
Liao Heng: I don’t think that is necessarily so. First, within universities, domestic compute resources are absolutely far more abundant than overseas — which is very surprising.
Zhang Xiaojun: Really?
Liao Heng: Of course. Look at Beijing and Shanghai — those national resource laboratories and the like all have scale of over ten thousand accelerator cards. That is an astonishing number. Abroad, at even the most famous university, a single department might have only 1,000 cards, or a few hundred. So in compute resources, China’s academia is absolutely far in the lead — abundant to an unimaginable degree. Of course, this is also thanks to certain senior leaders who, earlier in this process, recognized that compute had to be provided in order for academic research to be done well. I think this condition in China is unmatched. So as far as academia goes, China’s compute is very ample. As for industry, I have not seen excessive scarcity either — as long as a team is doing valuable work, they can still obtain a certain amount of compute. Is compute so over-abundant that it can be casually wasted? Certainly not.
Zhang Xiaojun: Do you think compute is a bottleneck for large models today?
Liao Heng: I don’t think it is the main bottleneck — no, it is one bottleneck among several. I can only say that there are some outstanding teams that achieved extremely high breakthroughs using very little compute. For example, the DeepSeek team mentioned earlier — their compute was quite limited, absolutely not on the scale of a hundred thousand or a million cards, but on the scale of a thousand cards.
And perhaps there is also Kunpeng (鲲鹏) — everyone puts the vast majority of their attention on AI compute, but general-purpose computing is also part of the infrastructure, along with its corresponding networking, optical modules, NICs, and SSDs. You could say we have assembled the full set of these ‘eight big items’ (八大件), which make up a complete data center. The most fundamental components, of course, are the CPU and the NPU or GPU; beyond that there is memory, the SSDs or related storage systems needed to hold the data, and then NICs and switches. The switch is also a component with a relatively high engineering threshold: the more ports a switch fans out, the stronger its interconnect capability — and it directly determines how many tiers the network ends up with. Say one switch can connect 512 things — then a single tier does the job. But if you need to connect 1,024 things, one tier cannot do it; you must have two. And once you have multiple tiers, there is latency — the extra latency between the two tiers and all sorts of extra costs — which double directly. And then there are the NICs too, right?
So when we build this system — especially these networking technologies — on one hand they need to be advanced, and on the other hand they need to interconnect. Interconnection means compatibility and inclusiveness, because within any system there is a great deal of old equipment; you cannot simply construct a brand-new world from scratch. So we have been trying to build — on the one hand, this LingQu (灵衢, UnifiedBus) system; and of course our Ethernet suite, built up over many years, is also very complete: from high-performance NICs to RoCE to switching equipment, in every lane we strive to be at least in the world’s first tier, with no generational gap. For example, if others are using 51.2T while I am still on 25.6T — at that point there is a generational gap. This factor matters quite a lot, because if we only made one single component, it would be very hard to form a complete system of one or two hundred thousand cards — there would inevitably be a large amount of legacy constraining your competitiveness. This is also the necessary condition for why we could step out and create an entirely new interconnect bus of our own — because as soon as you are doing interconnect, you are dealing with linking all the necessary components together, right? On one hand we have the complete Ethernet system, which ensures interworking with all legacy devices — and that interworking capability is not inferior; it is world-class.
But when we go to build this supernode (超节点) technology, it again involves those ‘eight big items’ I just mentioned — lacking any one of them creates a major defect; you could say every one is indispensable. This partly explains why the world has so many vendors — anyone who works in communications will tell you it is about standards: I have to go to IEEE, to IETF, and pass admission certification; even to do PCIe you have to go to compliance labs for interoperability testing, because it is a multi-vendor system — each component may come from a different vendor, and they have to fit together. So I think, especially when we started designing this LingQu (UnifiedBus) system in 2019 — on one hand we were forced into it, but even at that time we already knew that a technology system must be able to stand on its own, and the precondition for standing on your own is: for the things that need connecting, you must hold the full hand of cards, right? If there is even one component you have not assembled, you do not meet the condition — because that connection simply will not connect. So with these two factors combined — we happened to have fully independent capability in every key component, built to the point of being interchangeable with, on par with, world-class products — at that point we had what it takes to define a proprietary protocol of our own. And now we have in fact made the LingQu protocol free-license and published its spec, and we welcome vendors around the world to use it themselves. But from our own perspective, we had the ability to build a full-suite, complete system, and so we rather bravely took this step.
Zhang Xiaojun: Are there still any shortcomings?
Liao Heng: I think our biggest shortcoming is actually still at the software level. The hardware certainly has its deficiencies, as just discussed — we would like the specs to be bigger, right, with efforts over the coming years that may double every year; and memory — we hope to move faster ourselves so that it does not drag things back. But the biggest shortfall is still in the so-called software ecosystem. Our hope is that as our deployment volume keeps climbing, more and more people will have both the conditions and the necessity to join in contributing on this new hardware platform, working continuously on performance tuning. Because whenever a vendor or a customer deploys this system, they will necessarily have teams that need to pull their business up onto it. I think there may be a kind of threshold to cross here: once this population is perhaps double what it is now, what you might call the communal force — the strength of the group — will be large enough for the system to sustain itself in a virtuous, self-perpetuating way. So this is what we most look forward to.
Zhang Xiaojun: At the very beginning we talked about the ups and downs of the chip industry — you have been through its boom, then its sunset period, and then with the arrival of the AI wave it entered another boom. Now, in the application layer built on top of chips, the monopoly landscape is still unsettled. If one day they again form the kind of monopoly effect that those big American giants had back then, will the chip industry face another withering period? How do you see the future of the chip industry?
AI and the Chip Frontier
Liao Heng: I think… it is possible. It is possible. But there are certain factors here that keep piercing this — if that monopoly were a balloon, there are factors constantly puncturing the balloon. Because from Wall Street’s perspective: if I am an investor, and my entire pension is in OpenAI stock, or Anthropic stock, of course I want its returns maximized, right? It becoming the world’s sole provider of AGI models is most advantageous for me — from an investor’s standpoint. And if one company is not enough, what about two? Betting on two is not bad either, right — buy both companies’ stock at the same time.
Why do I think certain factors have kept this from happening so far? One reason — one important factor — is simply that AI technology has not yet reached its saturation point, its plateau; the curve has not yet entered its flattening phase. So anyone — even today’s leader — might be overtaken by another player six months, three months later, right? That possibility has kept alternating over at least the past few years. We once considered Llama the world’s best model, but now it clearly is — well, temporarily — not; maybe tomorrow it will stage a comeback, right? And within this, I have to say, the threshold of human IQ required is not actually that high. In other words, if you take a walk around Wudaokou (五道口), you will find perhaps 5,000 or even 10,000 students who can completely understand where the tricks in the latest models lie — and every day, in hands-on practice, they are running experiments at smaller scale, finding their next breakthrough point. So in other words, when something has not yet reached saturation, it is not easily monopolized — that is down to the technology itself.
The second important factor: I think behind this there may also be some people with real aspirations who do not want it — even if they had the capability, they do not want to become the world’s terminator; rather, they want, in a more inclusive way, to let more people join this field and make the next invention. In fact, among those I have been in contact with — at least the one or two leading teams in China — their entire vision of value is exactly this: they do not necessarily want to use their momentary lead to capture maximal short-term profit, but rather hope that, with this thing, smart people all over the world can keep chasing and overtaking one another, and so achieve the next breakthrough in capability. So I think there are two different value systems coexisting here — their visions are simply not singular.
And of course there are other factors as well. I believe that in an environment like China’s, there are a great many people — even our country itself — who do not want to see, absolutely do not want to see, all of AGI monopolized by a single American company. That could even produce a crisis for humanity, right? So: vision and values, plus drive — inner drive — plus China’s galaxy of brilliant talent and a very ample talent supply, with outstanding students emerging generation after generation — some of the most important inventions might even be made by an intern; that possibility has always existed. For this reason, I think open-source models are especially important — perhaps even the single most important needle puncturing this monopoly. On the hardware side, of course, we are making our own efforts too, right, to ensure it does not become a single, monolithic hardware system — and many of our industry peers are making this effort as well. So I think in the near term — at least for me — I remain full of hope, and believe this balloon is unlikely to —
Anyway, I’m firmly among those who hope to pierce that bubble. One more thing worth looking forward to: in the digital world — the virtual, digitized world — AI’s capabilities really do seem to be advancing very fast. Every aspect of my work now depends deeply on it, and I find that in many respects — almost every concrete task I give it — it does better than I do. That is what people mean by being very close to AGI in the digital world. Whatever your definition of AGI is, it is already useful enough, capable enough.
But for it to cross over into the physical world, there is still a considerable gulf. From our own seven years doing autonomous driving, my sense is — I already gave one reason just now for why legged robots may be hard to monetize for the time being, an energy-consumption reason, right? Maybe trailing a power cable solves that problem, but trailing a cable then restricts its working range. The bigger gulf, though, is that physical AI’s models have not yet reached their ChatGPT moment. That gap still needs — maybe it is the very next moment, maybe it is two or three years out — a major model breakthrough before physical AI can cross its zero-to-one moment.
And that domain will itself give rise to a fairly enormous, promising industry. Because once something is physical, it has to have a body — it must be a machine, right? And machines are inherently diverse. So what you can expect is that, as a result, once you enter the physical world, monopoly becomes much harder. Look at the car industry: it has been developing for over a hundred years, and there are still so many different brands, in different countries and regions, and new brands are still being born. Even people being fatter or thinner might call for two different sizes of car, right? So that is where the diversity lies. And what I find most promising here is China — I am very bullish on it.
I once heard that Buffett may have said something like “nobody wins shorting America” — meaning if you short America you are bound to lose. We don’t actually want to short America; I only want to — I want to long China. Because I think China itself, in terms of its conditions across the board, is very promising; it should be able to climb a very big step up. Moreover — including that interview you mentioned just now — many of these interviews seem to look at the question from a kill-or-be-killed perspective: if China wins, America loses; or that a Chinese superintelligence, a model of comparable capability, would be used as a weapon to attack American networks. I think that is a truly ridiculous perspective.
That perspective is basically a Satanic perspective. Because if I had such a good model, why would I use it to attack you? Why not use it to live my own life a bit better? Why not improve my own economy, my own healthcare, all those things — improving people’s lives, as they say — instead of using it to attack American networks? I think that very notion is a deeply wrong-headed motive.
Zhang Xiaojun: Is that a difference in values?
Liao Heng: I don’t know whether it’s a difference in values so much as a kind of pirate culture — they are accustomed to viewing the world through that lens. Even when China is relatively strong, it does not necessarily need to go and make others weak.
Zhang Xiaojun: Over the foreseeable next five or even ten years, do you think the global chip industry landscape will change? Will it go through some degree of reshuffling?
Liao Heng: The first question is whether it will. The factor with the biggest influence on this is whether we go back to the state before the US-China tech war and trade war — because that has very real short-term effects. If it continues another ten years, I think it will definitely split into two — even if there is no more serious conflict, everyone will need, simply in order to survive, their own complete manufacturing capability. Because chips, for modern life, more than oil — they have almost risen to the same level as water and air, or close to it. In terms of economic scale, it is certainly a bigger industry than oil. You can imagine: coming home to no rice cooker, no refrigerator, nothing at all with a power cord — that is unimaginable, isn’t it?
So I think it is a necessity. And that necessity leads to what you asked — if we’re asking whether there will be big changes, the first important macro factor is the impact caused by these factors between nations. I’m no expert in that area, so it’s hard for me to predict what that impact will be, but there will certainly be an impact. Only after that comes what you just raised — whether there will be monopoly, whether the withering-out phase I spoke of earlier will take shape. My answer is: not that fast. Maybe in ten years — hard to say, right? But right now, because it changes every day, you can’t even say who the leading vendor is in capability. So maybe you lead by three or six months, but in all likelihood someone else will reach a similar level — so the conditions for monopoly don’t yet exist.
Zhang Xiaojun: Looking at the floors above — the upper floors of this 18-layer pagoda — what would you want to say to these model companies or AI application companies? What expectations or outlook do you have for them?
Liao Heng: First of all, I think, number one: the best-endowed doesn’t necessarily win. It’s like when we all go to the same university with many classmates — you find that the classmate from the wealthiest family is not necessarily the one who ends up far ahead of everyone. Not that they’ll do badly, but they may just turn out ordinary. So the organization with the most money and the richest resources won’t necessarily — take America as the example: America’s top model companies are not Microsoft, not Google, not Meta, not Amazon. Why? They have plenty of money; if it’s GPUs you want, they have countless cards; their capacity to invest is strong; their research teams are large and formidable, right? Because — it’s a new thing, it’s not an old thing.
So I think a large, established company’s advantage lies in the enormous edge it has built in its own domain — otherwise it wouldn’t have become a giant, right? But the point is precisely that AI is a new thing. Whether in applications or in business models — the most valuable applications, the most valuable business models — they have yet to be invented, or are being invented right now. Look, I gave the AltaVista example earlier: Digital [DEC] was a giant in its day, and then it sold itself; AltaVista fetched a good price, but nobody ever figured out — such a great technology. When we were graduate students we thought it was wonderful: I could search for papers without going to the library. After it appeared I almost never set foot in the library again — perhaps only to photocopy one of the papers I had found, because back then PDF downloads were mostly not yet available.
So we have to look at the classmates upstairs — admittedly, each of those layers has countless giants in it. I expressed my hope just now: of course, if the giants can play to their strengths, provide really good services and keep improving them — for instance, I would love for AI coding to be a service provided by some Chinese giant, so I wouldn’t have to go to such lengths to make do with overseas ones, right? But I am also deeply aware that it won’t necessarily be them, because they may not want to be in this business. And of course I would very much like the best coding model to be Huawei’s — maybe I could just call the Huawei Cloud service directly. But every company has its own current state, its own culture, its own — the system it has already built. It’s like an immune system: if a cat’s cell suddenly appeared in my body, the immune system would clear it out. That is why the famous book we mentioned earlier, The Innovator’s Dilemma, actually describes exactly this problem. A mature organization usually cannot accommodate something new. But it depends on the organization’s own capacity for self-renewal — or on whether its founder has that vision, right?
The second point is that this vision matters enormously. If you say, I want to monetize this year and rapidly achieve monopoly, that is one approach. If you have a higher vision — say, by next year or in three years, there will be no more hard-to-cure diseases in the world, no more hard-to-debug problems, no more cybersecurity problems — that is a different vision, and it will drive your team and organization to push in different directions. But look at Anthropic — OpenAI created a chatbot, whereas Anthropic’s main revenue comes from coding, right, delivered through API calls. At that point these two businesses have nothing in common, don’t you think? Who they sell to, the different customers, how they charge — from a business-model standpoint there is no similarity at all. So I say: it’s a new business.
And is this model even the best model? Should you try to control the entry point to programming — say, control an IDE like VS Code? Or do you not need to — you sit in the back and just provide an API? You see, different people — even truly great figures — will throw themselves into ferocious competition based on some fantasy. For example, when the internet first took off, some people may still remember the battle between Microsoft’s IE and Netscape. That fight went on for years; in the end even the US Department of Justice intervened, right, and then Bill Gates retired. To us today that’s just a story. But when you look back — was that competition really meaningful? Today Windows doesn’t even ship IE; the default browser on Windows, the Edge browser, is Google’s Chrome browser with a shell around it. So the battlefield you once thought supremely important and must-win — after you won it, you discovered it’s not worth it, it’s nothing; it never brought you that fantastic result, right? Microsoft never became the hegemon of the internet, did it? Instead a mass of new hegemons appeared, and each of them found its own odd track. Nobody expected China to burst onto the scene — Alibaba, Taobao, Alipay, WeChat and all that — it was all new; they didn’t compete with the incumbents on the track the incumbents had defined, did they?
So I think AI’s application layer may be like that too. And for this application layer — in the digital world, the single most important threshold I think everyone needs to cross is this: AI is already extremely capable, very smart, but what it lacks is your personal context. If it does 95% of my tasks better than I do, the one thing for which it has not replaced me is that it has not been brought into my life context or my work context. So I have become a machine that feeds it prompts. Maybe I am inferior to it in every problem-solving capacity, but there is one thing for which it still cannot do without me: it did not attend this conversation between you and me, did not attend my meeting this morning, does not know what KPIs my boss wrote for me, does not know what the most important project of my past six months — perhaps hundreds of meetings’ worth — has been about. When you cross that threshold, you necessarily need an entry point. And that entry point, I personally expect, is the next super, super app.
[4:00:00]
Of course, this entry point will immediately intrude on human boundaries. For example, if it were embedded directly in all my social apps — in WeChat, say — so that the AI could see the context of my every conversation, then it would know everything about all my non-work interactions. It would certainly violate my boundaries, right? On the other hand, it might then no longer need me to describe long-windedly what I want done — it could just go do it for me. Likewise, if it were brought into every piece of software I use for work — able to see my every email, sit in on every meeting with me, see what I use at work, the WeLink that Huawei uses — it would know my entire work context: an always-present companion. It’s entirely possible that — I might be made redundant, because the next meeting might not need me to attend at all; it does everything better than I do.
But at that point it intrudes on another kind of boundary. For one, I would lose my sense of security; for another, my company would say: you have breached information security — these confidential meeting details, right, how could you let a chatbot or agent know them? Might it blurt them out somewhere inappropriate? You see, there is still an enormous amount to be done at the application layer here. I have only described what I think digital AI lacks: the link of bringing it into a person’s context. Whoever solves that first, whoever finds a novel design — and that design may have nothing whatsoever to do with the model itself; it’s a better app — that app may become a super, super, super application. Because if it knows everything you know, and it is more capable than you — then what?
Zhang Xiaojun: You mentioned the innovator’s dilemma just now — but this chip work of yours, within the Huawei system, should also count as a new thing?
Liao Heng: We are not a new thing — Huawei has existed for a long time, since well before I joined. But its dilemma lies in: who is going to define the future? Because this so-called future is a vision — what am I going to do, and how should I do it. In any organization there is a kind of — unless it is your own, unless you are the founder, unless you have every resource, right. Whenever it is not yours — when we are employees, or contributors — we have no choice but to persuade others: OK, here is the future, here is how you’re gonna do it. But often at that moment — for instance, the idea I described earlier of “living in 2030” — to another key decision-maker it may sound like utter delusion. So this requires a great deal of consensus-building. Call it a dilemma if you like — it is unavoidable, so we shouldn’t over-dramatize it either. In other words, in human organizational society, cooperation with one another is inevitable.
Zhang Xiaojun: If today were 2030, what do you see? You said you live in 2030.
Liao Heng: What I see is that we have sufficient manufacturing capacity; I see that perhaps digital AGI has become even better; I see that perhaps we are closer to having a machine that can sweep the floor for me, right. And then I would start to worry about what I myself should do.
Zhang Xiaojun: You have grown more confident over these years, is that right?
Liao Heng: The reason for more confidence is this: when you have faced difficulties ten times as great and found that in the end there was always a way to solve them, you naturally gain more confidence. Having faced those challenges and ultimately crossed those thresholds, it no longer feels so hard.
And there is one topic from earlier that I think deserves one more remark: optical communications. As we discussed, supernodes need to place chips farther apart, and therefore they need to use light. Here, once again, you see that we must respect physical boundaries. We actually built an optical module called Hi1 — it is a 7.2T optical module, whereas the optical modules on the market today are all 800G, that is, 0.8T. So straight away: what is 7.2 divided by 0.8? About 9 times.
Then we immediately face a very interesting question. Once everyone realizes chips have this much bandwidth that needs to be carried outside — with today’s optical modules, picture it: the machine is about the size of this table, your chip sits on a board, and the optical modules sit over on the faceplate side, connector after connector. From the chip to each connector runs a length of electrical cable — cables of maybe tens of centimeters, the longest perhaps 50 centimeters. Imagine it: it’s like an octopus, countless cables flying out of the chip over to a panel lined with optical modules. That is today’s prevailing industry model. The 7.2T optical module we just talked about, by contrast, is placed directly beside the chip. So the connection between chip and optical module becomes a distance of maybe just one to five centimeters — after all, the chip is only about that big. That eliminates that stretch of electrical cabling. This is what we — what the industry — calls NPO: Near Package Optics.
But when we set out to solve this problem, we very firmly chose NPO — we did not take the further step of putting the optics into the package, did not put them inside the chip; we put them beside the chip. And that question arises immediately. If you ever get the chance to interview other people in the optical communications industry, you will find this is a topic countless people have argued over for many years: beside, inside, or outside? What is very interesting is that, for my part, I didn’t need to go through the debate — I directly chose beside. In 2026, this 7.2T optical module absolutely had to go beside the chip — neither outside, nor inside the chip.
Why? Because of the second problem: wherever there is light, there must be a laser — what is the light source? And lasers fail; in a relatively hot operating environment they have a certain probability of failing. Once one fails, you have to solve the problem of how to repair it. And then there is a second question: a great many people, even having chosen to put the optics inside the chip or beside it, then firmly insisted on putting the light source outside. Because once it is outside, if it breaks I just unplug it — pull one out, plug a new one in, swap it — and that solves the failure problem. Whereas we firmly chose: absolutely do not put it outside, it must go inside. And not only inside — I put in two of them: if one fails, I still have another that keeps working, which lowers the failure rate. You see all this — and now you’ll start to doubt me: you, someone from a software background, an algorithms background, who then went and “pretended” to do chips — what qualifies you now to make pronouncements like this, that only this way is best? Well, this topic, once you get into it — it was actually quite a long process.
And it entirely bears out the philosophy I described earlier about how an individual grows within an industry. I mentioned two points: first, open up the funnel — look around, upstairs and downstairs, 360 degrees in every direction, and learn more. So around 2008 or 2009, I had already realized that chips would have to — because their IO, their electrical wires, can only transmit over such short distances — go optical some day. At the time we thought it would have to go optical at 56G, but in fact it only went optical at 224G — off by a factor of four. But even back then I realized we had to bring optics right up to the edge of the chip, or inside it.
With that question in mind, I set out to learn: how lasers work, how to design one, how to modulate the signal, how to receive the signal. And for this — because at the time the optics industry and the chip industry were two completely separate, isolated industries — in the first year I persuaded the company and signed up for a summer school. It happened that a university in the Netherlands was running one — Europe still had something of an edge in optical communications back then — saying anyone could come and learn how to design photonic chips. So I went during the summer holidays, for one week or maybe two. And I found that, OK, I now had a basic understanding of how optics works. That understanding actually became the basis for many important choices I made later.
I immediately realized that every optical device depends on one basic physical quantity, called the refractive index. For any material light passes through, the refractive index represents the speed of light in that material. In a vacuum, light travels at 300,000 kilometers per second, but in glass it is one speed, and in another material — silicon nitride — it is another speed. The second basic fact I learned is that the refractive index is changed by temperature, and no material escapes this: raise the temperature by one degree and the refractive index changes. Third, I learned that lasers fail. Fourth, I learned that for every laser, the light has to be coupled. With electrical wires, “coupling” just means, say, taking a wire and soldering it with an iron — it’s just a discontinuous point, and the wire is connected. But in optics, every single coupling may lose a few tenths of a dB of energy, because when light crosses from one material into another, it loses energy.
You see, these all look like they basically belong to high-school physics — basic facts that students who did particularly well might already know. But you find, surprisingly, that some people who have worked in the industry for decades have forgotten this knowledge — they learned plenty of other things instead. But this knowledge formed my basic view of the chip. First: why shouldn’t you immediately put the optics inside the package (封装)? Because inside that package it is extremely hot — I have a 1,000-watt NPU in there generating enormous amounts of heat. What does 1,000 watts mean? A household electric stove, a fairly high-powered one, might be 2,000 watts, right? So it is very hot. Put anything next to something that hot and it is bound to heat up, and the refractive index will change enormously. So putting it inside creates all sorts of problems caused by refractive-index shifts — best to avoid it.
The second problem: at high temperature, lasers fail. Third — I mentioned another reason earlier — I definitely did not want a single light source placed outside. Because if I have 72 optical channels and I want one light source to feed them, doesn’t that source need 72 times the energy? Seventy-two times the energy produces a hot spot. That is, an extremely high-power beam shining onto one point, which then also has to be split into 72 channels — that one point will inevitably fail. It’s that simple. With that much power — 72 times the power passing through a single point — that point forms a localized hot spot, and the hot spot will burn out your material. Or if glue or dust lands there, it instantly chars black, and it’s no longer an optical communications device — it’s become a cutting machine, right?
So we — perhaps others would do it differently — you see, all of these important choices I’ve just described were intuitive choices. I didn’t run experiments, and I didn’t have some hundred experts verify afterwards whether they were right. I didn’t need verification; I simply struck those options out, because I knew they were bad — so don’t do them. I chose what was easy to do and could be done well. And of course, our current 7.2T module has also turned out very well; in the next generation it may become the main way of interconnecting — it’s cheap, the bandwidth is high, the size is small, and it gets rid of all those extremely troublesome electrical wires. You see, this kind of thinking is purely intuition-driven, purely vision-driven; it doesn’t require discussion with countless people, and yet it may determine whether our system can be spread out across a larger physical space. These ideas all sound very simple when spoken aloud — but someone has to raise the question in the first place.
Zhang Xiaojun: After all these years, do you still feel strong passion for the chip industry?
An Engineer’s Story
Liao Heng: My passion comes more from necessity. Put it this way: if you’re walking down the street and someone — even a stranger — suddenly collapses, you’d feel you ought to go and ask whether something is wrong with them, right? What we see now is that this kind of need is very easy to perceive: we need compute, we need cars that can drive themselves, and so on — these needs are visible everywhere. We need our own smartphones; it can’t be that in all of China there are only iPhones, right? These needs are the proverbial “necessity is the mother of invention” — my passion comes mostly from that. But of course there’s also a negative driving force: when a person isn’t doing anything — or take myself — whenever I’m not doing something particularly good, the odds are I’ll go and do something bad.
Zhang Xiaojun: And what is the mission now? A personal mission, a corporate mission, or something else?
Liao Heng: Huawei’s mission is written out very clearly: to bring digital life to every person, every enterprise, every society. My personal mission is just the feeling that, in a finite life, I should leave behind as many better things as possible — or rather, that what I build should outweigh what I destroy. Because in life you bring nothing when you’re born and take nothing when you die; there’s only the process — and I hope to be of some use in it, of use to others.
You also wrote down two questions for me: one, where exactly is the AI Silicon Valley; and two, how do the ideals of AGI differ from the ideals of Wall Street. I actually already touched on both topics along the way. On Silicon Valley — I mean, where is AI’s Silicon Valley — I’d guess it’s probably in Wudaokou (五道口), (the university district in Beijing near Tsinghua), or else somewhere in Shanghai — Qiantan, or is it called Houhai or something, Binjiang? — over by Binjiang, those places in Xuhui district, the Shanghai Innovation Institute. Why do I make comments like this? Because my own child is over in Wudaokou, still in third year, and we also often go there for exchanges and discussions. I think if you sit down for a meal in any restaurant or café over there, you’ll find at the next table two people in a heated argument about why yesterday’s loss diverged, or how exactly to make reinforcement learning work a bit better. The atmosphere and the people are there.
And why not in Silicon Valley? As I said earlier — it’s not that Silicon Valley no longer has the cohesive pull to attract talent; it’s that these companies’ own desire for monopoly restricts the exchange. An Anthropic employee probably won’t discuss the latest reinforcement-learning techniques with some Stanford undergraduate, because they want to turn that into their competitive advantage — this so-called monopoly is in the process of forming. So I think China has a more abundant supply of talent, with more people actively and energetically working on all kinds of related problems — and for now China is much, much further away from that monopoly. So in this brief window — maybe these three years, maybe five — we are in a period of brilliant stars, a hundred flowers blooming, extreme prosperity, and extraordinarily lively exchange of ideas.
Zhang Xiaojun: So the formation of this open-source culture among Chinese companies — model companies, chip companies, companies of any kind — may have a more far-reaching impact than anything we can imagine right now?
Liao Heng: Yes, that’s how I see it. Or rather, I hope this open-source culture can last a bit longer.
Zhang Xiaojun: Last a bit longer.
Liao Heng: Yes. Because there’s no guarantee at all that the work people are willing to open-source now will still continue next year — so I hope they keep it going a bit longer. If it’s something the founder decides personally, it may last quite a while; if it’s just some department of some company, they may not even have the right to make that choice, right?
Zhang Xiaojun: Right — then they might go closed-source faster.
Liao Heng: Yes.
Zhang Xiaojun: Because it stays inside the company — you don’t know whether it will go closed.
Liao Heng: Right.
Zhang Xiaojun: With more challenges, it might.
Liao Heng: Yes, that’s right.
Zhang Xiaojun: I have a few final quick-fire questions. A favorite food of yours, from anywhere in the world?
Liao Heng: Food… These days I’ve been eating so much coarse grain that I no longer know what I like — or rather, I haven’t thought about what I like in a very long time.
Zhang Xiaojun: A little-known piece of knowledge that everyone ought to know? Actually, you already gave one earlier, so never mind. Based on all the books you’ve read, recommend a few.
Liao Heng: Let me recommend a few books. The first — we’ve recommended it internally at Huawei too — is called The Idea Factory; it should be available on JD.com, though I don’t know if the Chinese edition is titled “Sixiang Gongchang” (Thought Factory). It’s about the history of Bell Labs. The period that resonates most with us, and holds particular lessons, is 1940 to 1945. Because the era it describes was one in which America’s economy, its GDP, was already highly developed, but its science and technology still lagged far behind Europe’s old academic meccas — Britain, Germany. So in those days the top professors all had to have studied abroad — only a Cambridge graduate, or someone from Göttingen, or from some German university, could be a top professor.
But in those years, 1940 to ‘45, Bell Labs produced some fantastic work. Bell Labs sits right next to Princeton. Draw a circle with a 100-kilometer radius around it, and you’ll find that at that time, within those 100 — or 50 — kilometers, work of extraordinarily far-reaching influence was produced. First, in computing there was von Neumann at Princeton; then in information theory, Shannon; then for the transistor, Shockley, and also Bardeen, who invented the transistor, right? And the existence of all this was no accident — there was a kind of macro-historical, almost providential set of factors drawing them together. California had not yet risen as a technology center; the technology center of the day really was within that 100-kilometer radius around Princeton. That’s also why I say that when I was at Princeton I didn’t learn all that much — only later, after more experience, in my forties and fifties, did I suddenly come to appreciate, to realize, that the place really did have many of those elements of remarkable people and hallowed ground. I simply missed it, which is a bit of a pity.
And why is this body of work so instructive? Because first, it involves semiconductors. Then Shannon’s work is the foundation of every communication system; von Neumann’s work was precisely on the computing side; and Turing himself was also a Princeton student. So this whole series of works, compressed into a mere five years, actually mirrors and reinforces the big macro factors of our own current five years — maybe our current ten years. Because the Turing machine represents the deepest underlying theory of symbolism, while today’s AI represents connectionism — and connectionism is right in the middle of an explosion, isn’t it? Second, semiconductors — we’ve already talked a great deal about semiconductors; China is precisely in a state of extreme difficulty followed by rebirth, breaking out of the cocoon. And third, communications — China is no longer behind there. But communications and computing are, as I said earlier, one body with two threads — two lines entangled together through the entire industry’s development. So I think these factors — together with China’s real economy, which is already extremely strong, even unmatched in the world — mean that our science and technology must still leap from follower status to a point where, in certain areas, even the world’s most important inventions and creations are born in — perhaps somewhere in Shanghai, perhaps somewhere in Wudaokou, perhaps some combination of the two. I think in this era, this is quietly happening. It’s just separated in time and space — there may be an 80-year gap — but that 80-year gap marks the opening of China’s era.
So I think, regardless of whether there is a so-called trade war or whatever other factors, when these elements combine, we seem to have all the right ingredients — it happens to be the right time, right place, right ingredients — and a chemical reaction is bound to occur. So that’s one book worth recommending. And of course The Innovator’s Dilemma, which I mentioned earlier — that one is probably better suited to business leaders, right: realizing that your success will not keep on succeeding forever, or rather that the very factors behind your success are quite likely to be the chief obstacle behind your next failure.
Of course, there’s another book I’ve recommended and given to our colleagues, called The Rules of Work. It’s probably also sold on JD.com — I don’t know what the Chinese edition is titled. It’s actually extremely simple, a little booklet, about how young people should appropriately view the work environment. Because the vast majority of universities and schools will not teach you these things. The most essential point is this: in a work environment, unless you are purely on your own — maybe you’re a mathematician who can just shut yourself in a room and never interact with anyone — as soon as there is more than one person, there are relatively complex person-to-person interactions. So you need a basic understanding of this, a process of reaching consensus, and a way of operating within the rules — neither going to the extreme of revolution and wrecking the organization, nor letting yourself feel excessively wronged. Anyway, I’ve given this book to a great many colleagues; it helped resolve things for them and brought them a certain relief and comfort. So I think The Rules of Work is perhaps worth a read, especially for younger listeners.
That’s about it. Oh — while preparing, I also remembered the book Edison. First of all, Edison is a super-celebrity. Go back 100 years, and Edison was the Elon Musk of that era, or the Steve Jobs of that era. I think these few people may be heaven’s repeated resurrections — for all we know, the same soul reborn again and again. Or rather, I think they share certain common qualities. As for why I recommend this Edison biography — you can go to gutenberg.org, where there’s a free e-book, because I couldn’t find a print edition, it’s too old — it was written by someone who worked at his side, so it has a strong sense of authenticity. With Edison — people, especially very smart students who did well academically, are often confused: they assume Edison and Einstein are the same kind of figure, the brilliant giants of human history, the great stars in the minds of technically-minded people. In fact they are two completely different types. Edison may never even have attended middle school. He came from a very poor family, dropped out perhaps in primary school, maybe middle school, and then went off to sell newspapers on trains.
But he had an innate engineer’s quality. And this engineer’s quality really comes down to a few basic points. First: do useful things — never make things that are flashy and showy but without substance; you must invent useful things. Second: respect reality — you must first know which problems are worth solving, and not choose to waste time solving meaningless problems. Of course, the book has countless examples of what he invented, and to invent these things he basically used clumsy, brute-force methods — he was not a super-smart person. The so-called clumsy method was: OK, first — perhaps what was extraordinary about him was that he first identified what humanity needed. Hence the lightbulb: there was no electric light, so invent a lightbulb, right? There were no movies, so invent a way to record pictures of moving things — hence film; indeed even the first movie was shot by Edison. There are lots of interesting stories in there. I believe technically-minded readers, students with an engineering bent, will find plenty of meaningful references in it.
There’s one more book, which I read when I had just entered the industry, called The Soul of a New Machine — “Xin Jiqi de Linghun” in Chinese. It’s about — in my story earlier I mentioned a company called Digital, DEC, Digital Equipment — a company of the same era, the minicomputer era, called Data General. The Soul of a New Machine tells how a group of engineers, in perhaps the late 1970s, went about designing a new generation of machine. Why this book? It’s probably of most reference value to those of us who genuinely work on computers. It gives a very real, almost day-to-day account: how each person took part in the project, what difficulties they ran into, how someone was in a bad mood one day and how it got resolved. I even had the good fortune to meet one of the interns described in the book — back then I think he was called Bob. By the time I met him, he was already an EMC fellow, a white-haired old man. So what I want to say is: this book tells the story of engineers, and especially the story of an engineering collective — and that collective’s story is passed down from generation to generation; it keeps repeating. Maybe you were building a minicomputer; now we’re building a chip, or an AI supercomputer, or the next model, or the next promising super-app, right?
It was still enlightening for me, because you see the people who came before you — everything you have been through, the people before you went through as well. And he actually wrote it all down, which is notable especially because engineers are, frankly, a super boring group, right? Mostly introverted, not good at talking, and least of all inclined to write books or memoirs about themselves. So this is an unusual kind of book. What it describes is the real, almost first-person experience of an engineer. So I think for the engineering community, it is well worth a read.
Zhang Xiaojun: When you were talking about engineers passing things down generation to generation, I was thinking: everyone now says software engineers are about to be replaced by AI coding — if that’s the case, maybe people could pivot toward hardware.
Liao Heng: Not necessarily. Since we’re on the topic of coding, I think there are a few things. First, I don’t think you need to worry too much about being replaced. Or put it this way: as long as you have ideas of your own and are unwilling to be easily replaced, you will inevitably think up and invent new needs, and go build the next, more interesting thing — or, whether for your company or for yourself, define something you previously couldn’t do. Now suddenly, with AI behind you, you might have the development capacity of ten people, even a hundred people — so couldn’t you go build something completely new, something that has never existed before? Maybe you won’t earn a higher salary, but I think the thing you just mentioned, the thing so many people worry about — what you should really worry about is whether you’ll just “lie flat” (躺平), or whether you have an inner drive (自驱力) to create more useful things, services, or products. So as I said earlier — at least within our own small organization — I hope everyone has that desire to push themselves forward, rather than only doing something because I told you to do it. Based on our understanding today, what is the key, important bet? Betting on China — and that does not mean I stand for betting against America.
第二部分:中文结构化解读
本节为编者(非访谈者)对全文的结构化梳理与研判,不代表廖恒或华为的立场。
一、这篇文章为什么值得全文转载
- 人物稀缺性:廖恒(Liao Heng)是华为 Fellow、华为半导体首席科学家,自 2020 年华为「至暗时刻」以来,极少有华为高层如此公开、系统地谈论昇腾(Ascend)如何从谷底一步步走出来。这篇 4.5 小时对话是难得的「体系视角」而非单点技术秀。
- 来源链条:访谈由张小珺商业访谈录(语言即世界工作室)于 2026-07-25 在 Bilibili 首发(中文原声);本文文本为 Hamish Low 用 claude code(Whisper large-v3 转写 + 翻译校订)整理的英文全译,发布于 Cambrian(Substack),2026-08-04。原文即 AI 转写翻译,个别术语/数字可能有误,阅读时宜保留判断。
- 价值:用一条「纵(历史)+ 横(供应链)」的框架,把全球半导体 40 年、AI 算力硬件的演化与中国路径串成一张图。
二、核心论点拆解
1. 垄断会扼杀创新 —— 半导体为何曾是「夕阳产业」
PC/互联网时代的赢家(Google、Meta、Windows、Apple)在 3–4 年内建立近乎垄断的优势,并凭资本、基础设施、用户习惯锁定需求与技术方向。当单一客户占据绝大多数采购量,供应商便被压价、被锁定路线。结果:2005–2015 整整十年,硅谷 VC 对芯片初创零投入,Broadcom 式「并购—裁员—提效」成为行业收割模型。廖恒的逆向洞见:正是上一波的极端成功,造成了下一层的枯萎。
2. 创业者的窘境(Innovator’s Dilemma)
首创新垄断者难以自我革命。今天 AI 领域取得最大突破的,恰恰不是上一代巨头——「往往不是最有钱的孩子最成功」。
3. 代工模式的经济学根源,与「大前提」的崩塌
晶圆厂资本开支随节点指数级膨胀(掩模层数 10→100、单机台 10 万→1 亿美元),设计公司无力自建——于是台积电开创的 fabless 模式成为公共基础设施。「real men have fabs」最终也被证伪(AMD 自身拆分制造)。但这一模式的大前提是「世界是平的」、全球分工最优;2019 年后地缘竞争彻底推翻了该前提,加上 DRAM 半年涨十倍这类剧变,垂直分工无法应对超高速变化。
4. 摩尔定律的真相,与 Tau 定律
- 经济维度:约 16/7nm 起,单晶体管成本已不再下降、反而上升——反通胀红利消失。
- 性能维度:增益极小。
- 能效维度:仍是唯一真实获益处(bucket 变小,每次翻转耗能更低)。
- 廖恒主张改用「原子数」而非「纳米」度量(已逼近单原子尺度),并引出华为 2026 年 5 月末发布的 Tau 定律(时间尺度:我要电路更快)与 Carnot/Joule(能量尺度:同样任务更省电)。换个考题,答案就变——这是全文的方法论底色。
5. 十八层宝塔 vs 黄仁勋五层蛋糕:协同设计(co-design)是核心
芯片位于第 7 层,承上(编译器/并行划分/RL/KV cache/压缩/稀疏化/新数据格式/应用)启下(器件/工艺/材料/矿产,即 5–6 层「地下室」)。真正的稀缺能力是跨层——像一根线把珍珠串成项链。廖恒自嘲是「糊涂(十八层糊涂)」的跨层观察者。
6. Cube/Vector 算力比:8:1 与 32:1 的玄机
Blackwell 的 Tensor Core(Cube)相对 Vector 为 32:1,昇腾为 8:1。高冗余(32:1)允许粗放用算力;低冗余(8:1)逼迫算法做稀疏化/压缩——DeepSeek 的 MoE 稀疏激活(省 32 倍算力与访存)、长序列稀疏注意力即典型。结论:华为芯片规格约为 Blackwell 1/4,跑 DeepSeek 却「够用」,正是协同设计的胜利;梁文锋主动选择更高复杂度、提前解决预见的瓶颈,是「有远见者」的价值。
7. 中国半导体路径:910 → 950 与「分解方法论」
- 910:好日子,TSMC 最先进工艺 + 全球最优 EDA/IP。
- 2019 断供 → 950:国产工艺/工具/全栈,靠「分解 + 跨十八层协同设计」补位(堆叠、新数据格式、与 DeepSeek 等算法方共创)。
- 方法论:把看似不可能的问题分解为 100 个具体问题,再分解为 1000 个物理/化学/数学问题;一旦可触(tangible),100 个里能解 80 个,80 个里再有 10 个做到比别人好——以长处补短处。
8. 中国的相对优势
应用层极强(支付宝/微信支付全球独一份的便利)、算法人才(「全球 70% 杰出算法人才是中国人」的夸张表述背后是科举传统与上升通道)、能源(电力供应≈美国 3 倍、数据中心电价≈美国 1/4~1/5)。即便能效落后,「算力够用且便宜」——断供也绝不等于「无算力」。
三、与本站已有系列的衔接
- CPO / 光互联:访谈未直接谈光,但「地下室」第 5–6 层(工艺/封装/堆叠、互联下沉到 chiplet)正是本站 CPO 测试、先进封装主线的底层——廖恒所言「改变结构而非继续微缩」(如把两条平行导线改为正交以降低电容)与光互联/3D 堆叠解决单节点受限,是同一思路的实例。
- 先进封装:十八层宝塔的「缺 EUV、缺设备」困境,对应本站先进封装系列(AMD/高通 EMIB、Medha/Hamsa/Owl 拆分、CoWoS-L)。
- AI 硬件 / 算力:与本站 AI 硬件入门、AI 内存入门、推理芯片架构地图衔接;Cube/Vector 比与 DeepSeek 协同设计,是「国产算力够用论」的底层逻辑。
四、投资映射与风险提示
- 核心判断:「押注中国,但不是押注看空美国。」
- 对投资的几条含义:
- 应用层/算法层中国已具全球竞争力——关注真正具备跨层协同设计能力的系统厂,而非单点参数玩家。
- 设备/材料/工艺(「地下室」)是确定性补短板的长期赛道,但周期长、ROI 极差(廖恒亲证「晶圆厂 ROI 极差、回收期长」),需以长期资本视角而非主题炒作对待。
- 能效差距被能源成本与供给优势对冲——算力供给的安全边际高,是国产算力「够用论」的硬支撑。
- 警惕「垄断未形成」窗口期被证伪:若美国 hyperscaler + NV 形成事实标准并锁定生态,协同设计红利会收窄。
- 风险:纯 AI 转写翻译,术语/数字可能偏差;廖恒自陈是「悲观底色」的跨层观察者而非单点专家,其判断含主观框架,宜交叉验证。
五、一句话总结
廖恒给的不是一个技术答案,而是一套「分解问题 + 跨层协同 + 押注中国人才/能源/意志」的世界观。对投资与产业研究的启示是:不要在单点上与 NV/TSMC 比参数,而在「度量尺度(Tau/Carnot)」与「系统协同」上寻找非对称优势。