While the market agonizes over whether big tech's hundreds of billions in AI capital spending will ever pay off, hardware engineers on the front lines are focused on a far more fundamental issue: the economics of Moore's law have collapsed, and the true bottleneck in the compute race never was the chip itself. The challenge has shifted from designing faster processors to controlling the entire physical stack, from silicon all the way to the rack.
In a recent interview on the Frictionless Podcast with Logan Jastremski, "Mr. Bubble," a prominent chip designer and hardware analyst, dissected the core constraints of today's AI supply chain, the future direction of architecture, and where the ultimate value in AI investment will land. The conversation, ranging from transistor density to memory hierarchy and rack design to the strategic intent of frontier labs, converges on a single conclusion: the real obstacles have moved to advanced packaging, interconnect latency, and memory hierarchy.
From the viability of NVIDIA's roadmap and the memory race beyond HBM to divergent inference architectures and the real motives behind AI labs "pacing the frontier," Mr. Bubble offers several contrarian takes. His analysis suggests that instead of betting on rival frontier AI labs, investors would be better served by holding the foundational hardware and supply chain companies.
The Economic Collapse of Moore's Law and the Rise of Advanced Packaging
For years, the market has measured chip progress by Moore's law, the doubling of transistor density roughly every two years. But Mr. Bubble stresses that this logic has broken down economically. "The capital required to push these process nodes forward has ballooned to staggering levels, now potentially costing tens of billions of dollars. And you can't even double the transistor count anymore; you might only get a 15% to 20% density improvement." Because the cost of lithographic scaling has grown exponentially, he argues that "advanced packaging has become the new Moore's law."
The industry is no longer fixated on cramming more transistors into the same area but is instead "expanding the area" by connecting multiple chiplets so they operate like one. This trend is redistributing value across the supply chain. Mr. Bubble credits this insight for his early bullishness on Intel, noting their EMIB packaging technology was far ahead of Taiwan Semiconductor Manufacturing's offering at the time. However, he admits that "the real long was actually PCB board makers," because when NVIDIA was forced to downgrade its Rubin Ultra design from four chiplets in a single package to two separate packages connected via PCB, the circuit board manufacturers became the new bottleneck in the compute chain.
Using NVIDIA's roadmap as an example, Mr. Bubble exposes its dependency on external factors: "If you look at Nvidia's roadmap, it's not controlled by Nvidia; it's constrained by memory and packaging." He reveals that NVIDIA initially planned to put four dies in a single package for Rubin Ultra, but physical limits forced it to compromise (rumored to be two dies per package with two packages connected by PCB). "Now the PCB maker is the bottleneck for Nvidia to make these chips, which are centimeters apart, behave like a single chip."
"Every time I look at Nvidia's roadmap, I see 'impossible today, impossible today,'" he emphasizes. The Feynman generation is supposed to come with four chips by default and more HBM, things "we intuitively know are impossible today, and I'm curious how they plan to solve it." He pushes back on the idea that "Jensen can do anything," stating, "My response is: he is constrained by what SK hynix, Samsung, and TSMC can do, and maybe in the future, what Intel can do."
The Real Unit of Compute: The Entire Rack, Not the Chip
As large model parameters scale up, the center of gravity in compute is shifting from horizontal scaling (scale-out) to vertical scaling (scale-up), fundamentally reshaping data center architecture. "You can no longer design a chip in a vacuum," Mr. Bubble asserts. "These systems are no longer just one or two chips, or even a server; they are an entire rack, possibly multiple racks. In the future, whether for training or inference, the design focus must be the entire rack."
In the vertical scaling domain, interconnect throughput and latency determine a cluster's viability. Mr. Bubble contrasts two radically different approaches: NVIDIA solves high-bandwidth interconnect with its NVL72 system and massive NVSwitch network, which consumes physical space that could otherwise hold GPUs. In contrast, Alphabet's TPU team has delivered an impressive alternative with a scale-up domain of 9,600 chips, building routing logic directly into the chip itself to act as pseudo-routers.
"Based on your scale-up domain, you can change your chip design," he explains. With a low-latency, high-bandwidth domain and fewer intermediate switches, a chip can carry less memory, opting for "more chips instead of more memory." This explains why Google's TPUs have less HBM per card than NVIDIA yet can still handle hyperscale training. Mr. Bubble believes this system-level design capability will become a key competitive moat for AI hardware vendors.
On the inference side, "prefill-decode disaggregation" is becoming the dominant architectural paradigm, with Kimi's reference design deploying the two phases on separate nodes. "This is becoming ubiquitous because these two operations are fundamentally different things."
After HBM: The Real Memory Battle is Flash
On the AI inference side, the exploding context window—from 128K to a million tokens or more—is putting unprecedented pressure on memory capacity. The market assumes an endless need to stack expensive HBM, but Mr. Bubble proposes a pragmatic alternative. "At the end of the day, model weights should be in HBM. But if we don't use all the context at once, why should I put a ton of context data in HBM?"
He notes that it's user history and context, not just weights, that consume massive memory space during inference. The industry is moving toward "flash offload" to resolve the throughput-cost dilemma. NVIDIA has already introduced network-attached flash drive rack designs for these scenarios, and cheaper, higher-capacity HBF or even regular flash, paired with excellent flash controller design, can handle the massive context demand.
For the ultimate future of memory, Mr. Bubble offers a clear prediction: "If someone wants to know what I think the future of memory is, it's 3D DRAM." He suggests bonding DRAM directly on top of logic chips, leveraging proximity to simultaneously reduce latency, increase bandwidth, and lower power consumption—the next hardware revolution.
Own the Companies That Build the Railroad
Turning to the recent "pacing the frontier" narrative from frontier labs, Mr. Bubble questions the true motive: "Is it for us, not for themselves?" He points out that these labs haven't paused any data center construction or cut capital expenditures. "Those data centers are multi-year commitments. If they were sincere, they would cancel that spending. That's not going to happen."
On AI safety, he is not a believer in doomsday scenarios: "The cost of keeping these GPUs running is extremely high, so it going to take over the internet and break everything sounds like a waste of resources." He sees legal liability as a more realistic risk than a sci-fi runaway.
His real question about "pacing" is how much is about tackling high-value problems versus "regulatory capture." He gives examples of solving millennium problems like Navier-Stokes for "$15 million to make a million, fine, but that problem might not be that valuable," contrasting it with Terence Tao's math research that underpins LTE communications, which is genuinely worth billions.
His investment philosophy is unwavering: "I'm a firm believer in owning the hardware layer, owning the railroad, or owning the core components that make it all up." He asks a pointed question: "Do you want to invest in Anthropic and OpenAI—companies that will compete to the death and whose alpha will dissipate as people switch jobs? Or do you want to own the companies with 10- or 20-year track records and deep moats in developing indispensable hardware?"
On the "capex payback" question that worries Wall Street, Mr. Bubble has a simple economic retort: "I think we're going to have a fast takeoff. Just five to ten new drugs developed with AI will rationalize all the capital expenditure."
Chapter One: Introduction & Why Hardware is the Real Core Topic
Mr. Bubble: We're going to add more memory to these servers. The question is what kind of memory: HBM, DRAM, or flash? My guess is more flash. Everyone is talking about reducing the amount of HBM we need. At the end of the day, weights can live in HBM, but if we're not using all the context simultaneously, why have a ton of context sitting there? I am a firm believer in owning the hardware layer, or owning the railroad. So, do you want to invest in Anthropic and OpenAI, which will compete to the death and whose alpha will spread as people move around? Or do you want to own companies with 10-20 years of R&D history and moats in core hardware? They are advancing the frontier for us, not for themselves. I tend to be very bullish. We'll have a fast takeoff. I think just five to ten drugs will justify all the capex.
Logan Jastremski: Thank you for being on the podcast. I'm excited to have you. You have a remarkable talent for explaining technical concepts simply. I look forward to unpacking what you're seeing today and where things are going.
Mr. Bubble: Thank you. My "talent" for simplification is because I'm first and foremost a chip designer. I must simplify or get bogged down in details. Also, as a trader—or gambler, as I like to say—you only really need a high-level overview. People get lost in the weeds arguing about a SerDes bit error rate or a material saving a few watts of power. I say, guys, come on, what is this thing supposed to do? That's its economic value to the world. A little bit here or there doesn't matter.
Logan Jastremski: We have many nerdy debates in crypto too. Simplify, then add importance, as Elon likes to say. The simplest architecture usually wins. There's a lot of that in AI.
Mr. Bubble: I used to love philosophy, reading ancient stuff up to the Stoics. They had this concept about the essence of things, like Plato's theory of forms. What is the essence? What is the core? From that core, you can derive everything else. That's how I approach these topics—what's the main thing? Then we talk details and see if they matter.
Logan Jastremski: Let's start from zero to one with something basic like Moore's law and why it's ending. Then we'll jump into more interesting content.
Mr. Bubble: Before discussing Moore's law, you need to know what it says. It's about transistor density, doubling roughly every 18-24 months. This matters because transistors build digital logic, which comes from programs. The more transistors we can fit, and the smaller they get, the more work a computer can do. It's economically vital as well. You can rationalize capex, predict a return on investment, charge a premium, and gain market share. But what people don't realize is a branch of Moore's law—let's call it "Bubble's law"—is the cost of improving density. Early on, it was cheap. Now, you might have to spend hundreds of billions to not even double, but only get 15-20% density improvement. The economics have broken down. Only a few companies can afford the investment. If you don't invest, you fall behind, creating a vicious cycle where you have no customers, no money, and no time to iterate. The key is the money. To double transistors, from the 60s through the 90s, was almost free. We didn't have to think about it until the early 2000s.
Chapter Two: Chips from Zero to One
Logan Jastremski: At some point, we transitioned from single-core to multi-core performance. How did that change the dynamics?
Mr. Bubble: That was mid-2000s, maybe with the first dual-core Pentium. A core processes instructions. How fast depends on clock frequency, which has hit a ceiling. This is related to switching speeds. The world realized some instructions can run in parallel. Graphics is the obvious example—embarrassingly parallel workloads. Matrix operations are embarrassingly parallel. That's why multi-core systems rose. For certain workloads, it's a way to gain performance. If you have a single-threaded workload, a good programmer could run it on an old PC with a high clock. The Core i9-14900KS is the highest-clocked consumer processor, hitting around 6 GHz. No single core is faster today. For most, that's not important, but for things like network packet processing, where you're receiving data byte by byte, clock speed is everything.
Logan Jastremski: That gave way to NVIDIA and GPUs. Their matrix math maps well to AI. How did AI start scaling up?
Mr. Bubble: Why does NVIDIA exist and why are they good at this? Graphics processing is really rotating shapes—triangles and squares—based on matrix multiplication. That's the same math used in AI. The benefit with AI is you don't need high precision, whereas graphics does. For physical simulation, you'd use 64-bit floats. NVIDIA wasn't the only graphics company, but it was clear you needed a specialized processor because dual-core CPUs were nowhere near enough for a 60 FPS game. The interesting question is why NVIDIA dominated for so long. The answer is that they designed their GPUs to be flexible and programmable from the start. There's a famous story where Jensen says NVIDIA almost went bankrupt because they didn't support a new graphics API. They had to quickly tape out a new chip or be killed by Microsoft. Even then, they focused on parallel IO and virtualization. They emphasized programmability. Many companies that went fixed-function died. NVIDIA's moat remains deep. New models come out, and NVIDIA supports them in a day or two. They expanded beyond spec, went into scientific computing, and that was their springboard into AI. That was the right choice, and it's clearly paid off.
Logan Jastremski: You're known for your call on Intel. Can you explain why packaging is the new Moore's law?
Mr. Bubble: If we could shrink transistors forever, packaging wouldn't matter. But lithographic scaling has stalled. We might get 10-12% improvement for billions of dollars. A smarter approach is to add more chips together and connect them to act like one. This increases area. If you know anything about lithography, taping out at an advanced node is astronomically expensive. Doubling the area is a new scaling paradigm. Packaging is how people are now trying to expand transistor counts. I was focused on Intel because they tried to leapfrog TSMC by adopting high-NA EUV early and pushing packaging. It requires massive investment. It becomes a monopoly or duopoly. The only way to catch TSMC is to burn money. I had high confidence in their packaging tech, EMIB, which was far ahead of TSMC at the time. Now TSMC is moving in that direction. NVIDIA's roadmap depends on memory and packaging, not anything they control.
Chapter Three: Packaging as the New Moore's Law
Mr. Bubble: NVIDIA doesn't make or design these things. It's a H100 with one big chip, a B200 with two, Rubin with two, and Rubin Ultra was supposed to have four. Guess what happened. Rubin Ultra can't fit four chips in one package. They've had to downgrade multiple times. The rumor is two chips per package and two packages. They're connected via PCB. The real opportunity might be going long PCB makers, because they've become NVIDIA's bottleneck—making two chips, inches apart, behave like one isn't easy.
Logan Jastremski: The bottleneck is improving data throughput between chips.
Mr. Bubble: Throughput and signal integrity. As distance increases, signals lose integrity. We have power lines and repeaters for phone lines. On a PCB, signals need conversion, transmission, and error-free reception at high speed. Rubin Ultra hasn't shipped yet. They're having significant trouble. Feynman was supposed to have four chips and more HBM by default. All things we intuitively know are impossible today. Every time I look at NVIDIA's roadmap, I see "impossible today." People say, "It's NVIDIA, Jensen can do anything." Well, he's constrained by what SK hynix, Samsung, TSMC, and maybe Intel can do.
Logan Jastremski: We've scaled from chips to racks to hundreds of thousands of connected GPUs. How does scale-up work for large pre-training data centers?
Mr. Bubble: Why scale at all? You can't put 72 packages into one package, like the NVL72. But AI algorithms are highly parallel. One chip can handle part of a problem, another chip can handle another part. If it were single-threaded, scaling would be useless. In scale-up, latency and throughput matter, but latency is more critical. For pipeline or tensor parallelism, the data transferred isn't huge. NVIDIA's systems are for massive bandwidth with NVLINK—about 900 GB/s. This is more relevant for training. Other companies approach scale-up differently. At Hot Chips, the TPU team built a scale-up domain of 9,600 chips—an astonishing number. Their on-chip HBM is less than NVIDIA's. They don't do multi-die packaging. They specialize in scaling the interconnect network and routing traffic between chips. NVIDIA's NVL72 system uses four NVSwitch units, taking up rack space. The TPU team put routing logic directly in the chip, making them pseudo-routers. Scale-up is becoming more critical. You can't design a chip in a vacuum. These systems are a full rack, maybe multiple racks. Depending on your scale-up domain, you can change chip design. If you have low-latency, high-bandwidth scaling and less switching, you can reduce on-chip memory and use more chips instead. It's an art of balance.
Logan Jastremski: Disaggregation is interesting. The same hardware can be optimized for memory or compute.
Mr. Bubble: Yes. The industry is realizing that different operations should run on different silicon or cores. Kimi has separate nodes for prefill and decode. This is becoming the standard way to serve inference because they are fundamentally different operations. The future of AI training and inference is designing the entire rack system. We can't design a chip in a vacuum.
Chapter Four: Racks, Trays, and Compute Units
Logan Jastremski: Besides packaging and interconnect, what are other bottlenecks in training?
Mr. Bubble: NVIDIA chips are very strong for training—great floating point and bandwidth. There are fewer hardware optimizations left. You can offload things to flash or save computations for later, but that's algorithmic. It's getting harder to pack more processing power. If you want a crazy compute chip, you'd need a large die, a huge area, with massive multipliers—Cerebras style.
Logan Jastremski: Isn't that what Elon intended with Dojo?
Mr. Bubble: I don't know.
Logan Jastremski: They shut it down, so I guess not.
Mr. Bubble: I'm not surprised. In training, you have massive amounts of data—activations, states, weights, checkpoints. There are huge data problems. Offloading to cheaper storage might help, but there's less room for improvement. We rely on algorithmic people to figure this out rather than hardware.
Logan Jastremski: I ran an experiment looking at the top 10 models on OpenRouter. Even the smallest needs 2-4 H100s, each costing $20k-$30k. For GILM 5.2, a trillion-parameter model, you'd need 40-50 H100s just for the weights. What's the impact on memory? SRAM? HBM? Flash?
Mr. Bubble: You're numbers are roughly right. Even Llama 70B needs about 4 GPUs. What you store in memory matters. Weights are scaling, but there's an argument that it's flattening. GPT-6 reportedly uses a technique to loop the same weights, a transformer. If adding more weights worked, we'd need more memory. Research shows many weights are empty or sparse. Quantization and pruning work well. Speculative decoding is effective because not every weight contributes equally. But what takes up space? Context. The user's dialogue or previous conversations. Because these models generate tokens auto-regressively, they traverse the entire context. Context expands with weights, attention heads, and users. I'd argue context can be more important than weights in inference.
Logan Jastremski: The context window is fascinating. It's grown from 128K to a million. Where does it go? Ten million? One hundred million?
Mr. Bubble: Pure expansion is hard due to hardware limits. Even a million isn't a true million because training on that much data is difficult. But more context will come, offering more information about you. How do they get that information? Maybe through cached conversations retrieved later, or post-training on your context. If something is important enough, you burn it into your neurons. For less critical things, you search your files. We'll store a lot of data about people—think Alphabet and Meta. The benefit is avoiding prefill. If you already have the context, you can start decoding immediately.
Chapter Five: Memory Hierarchy: HBM, DRAM, and the Next Level
Mr. Bubble: GPU utilization is much lower and power consumption is less. Services charge differently for context hits. Kimi charges less if you have context set up. Context may not live in the physical hardware itself but in system design. NVIDIA has rack-level flash storage for this. For agentic work, context is key. Agents are just a framework, a pile of prompt injections. If you execute that prompt every time, it consumes GPU and your first token time is forever. So caching the agent framework is beneficial.
Logan Jastremski: I'm interested in maximizing decode throughput. SRAM has high throughput but limited capacity. HBM has more capacity but still isn't enough for the fastest speeds. You've been a proponent of high-bandwidth flash (HBF). Is the industry over-optimizing for throughput?
Mr. Bubble: The industry focuses on throughput because higher token throughput earns more money. If I constantly have to fetch context, it's a compact footprint. Perhaps you can get better throughput by being smarter about context. I joke that to change the world, you either cure cancer or invent a new type of memory. Flash has better bit density. If we can get competitive bandwidth, we could have more capacity at lower cost per bit. The reliability issue is solvable with a great flash controller designed for context, not for databases. Context doesn't always need perfect preservation. The problem with HBF is finding its market. If you have HBM, why use HBF? Does more GB help? Not necessarily. Decoding is about bandwidth. Density only helps store more context. If you're smart about moving data, you can store context on regular flash. HBF latency is 5 microseconds, similar to high-end AI-specific flash. Its bandwidth is lower than HBM but cheaper. The customer is someone who can't get HBM. I've heard from a couple of sources that the main HBF customers are the Chinese, who are very interested. This makes sense, as they can't manufacture high-end HBM. They might accept an inferior but cheaper product. HBF is early stage. We need standardization and to co-design the entire rack for AI.
Logan Jastremski: You prefer offloading to flash over HBM?
Mr. Bubble: I don't have a strong preference. It's better than HBM for some cases. NVIDIA thinks so. But I think we can do better. We need intelligent orchestration across all levels, including compression algorithms for context.
Logan Jastremski: Will we get from a million context to ten million? Driven more by algorithms or hardware?
Mr. Bubble: Inference might start to look like training with test-time compute. If we're training, we might not need as much flash. But we're at the frontier; we don't know the best tech. You'll offload some things to flash, but not everything. For enterprises, there's the logic of feeding a database to an LLM. But it's uncertain.
Chapter Six: Vertical vs. Horizontal Scaling
Logan Jastremski: Will the industry prioritize throughput for revenue or specialize in longer contexts?
Mr. Bubble: The industry is focused on throughput. These companies make money by serving more users. With an NVL72 rack, flash offloading, or algorithmic improvements, you could double or triple the number of users. That's more valuable to an AI lab. Ultra-fast decode models will exist, but the fundamental value is in serving three times the users on the same hardware, lowering cost per token. New capabilities might be monetized by enterprises doing post-training on their databases, like Thinking Machines.
Logan Jastremski: Is there a way to play this publicly? Long Cerebras? Or Etched?
Mr. Bubble: Long Etched would be a great investment, but they're not public yet. On the public market, it's hard to find a clear play. Many interesting AI companies are private. You can try to gain exposure to the hardware or commodity side. There are interesting companies like Weka, which provides software for flash offloading or storing data across clusters. There are opportunities in software and flash controllers. No public company's compute exceeds NVIDIA. If you believe context will be burned into weights, you need compute for post-training, so anyone providing low-cost compute becomes valuable.
Logan Jastremski: For ordinary people who can't access private markets, is the best way to play this through companies using agents?
Mr. Bubble: That's counter-intuitive. Maybe insurance companies. Agents will be like the internet or cloud. If you don't have an agent use case, you're not in the game. The big tech companies adopting AI are also the ones providing it.
Logan Jastremski: Are you interested in CXL to improve memory bandwidth?
Mr. Bubble: Yes. In decoding, throughput equals memory bandwidth. CXL allows memory to be added and appear like standard memory. It's useful for offloading context, like vLLM does to RAM. Building a direct way to do this adds value. The same hardware, new types of memory, and better code can get more value from the same accelerators.
Logan Jastremski: What about photonics for compute?
Mr. Bubble: I should write a blog post with the math. But look at the size of our AI models and matrices. Optical matrix multiplication happens in the analog domain. You'd need a stadium-sized space to get equivalent matrix dimensions. I'm not bullish on photonic computing. Maybe for IoT or embedded devices like Waymo's LiDAR. Some people are selling a dream. Let's see how long it takes.
Logan Jastremski: What excites you from an engineering standpoint?
Mr. Bubble: Memory. It's why I've always liked HBM. We'll see more innovation in memory. Optimizing memory locality is key. Jalapeno, OpenAI's chip, showed this. Memory channels are physically near compute, so keeping compute close to specific memory groups improves performance. 3D DRAM, integrating memory on top of compute, is the future. You'd have tiny interconnects, so spatial organization matters. This could provide more gains in throughput and power without better HBM.
Logan Jastremski: Is this about latency or throughput?
Mr. Bubble: Both. Better locality means lower latency. More memory channels mean more bandwidth. You can have just enough bandwidth for the specific block. But latency is much lower. People talk about bandwidth and latency but not *when* you need data. If a GPU operation takes 10 ms and I need data at 11 ms, does it need maximum bandwidth or latency? No. You need to schedule fetching. That's related to locality. 3D DRAM is the future.
Chapter Seven: AI Alignment
Mr. Bubble: There's a company called D Matrix that showcased stacking DRAM over logic at Hot Chips. Hot Chips was great. I moved to Palo Alto partly because of it. Last year, it felt like the Enlightenment era in London, full of legends. This year was packed with finance people. I found it interesting. I think 3D DRAM will be huge. It offers higher bandwidth and lower latency, though harder to design. On power, Jensen's point about 1 GW data centers is right. We can reduce power, but it might sacrifice token throughput or the number of users. The cost of data centers is not electricity; it's hardware. So if you start a new cloud company, power is irrelevant. The GPU cost is the issue.
Logan Jastremski: So memory is the biggest concern, given context lengths and more users?
Mr. Bubble: Yes. We'll add more memory to servers. The question is what type: HBM, DRAM, or flash? My guess is more flash. Everyone is talking about reducing HBM usage. Weights should be in HBM. If we're not using all context at once, why have it there?
Logan Jastremski: Any other interesting topics?
Mr. Bubble: I want to talk about "pacing the frontier." I'm curious what it actually means. It hasn't been clearly defined. They solve millennium problems for $15 million to make a million. But is it valuable? The Navier-Stokes solution might only help CFD simulations. There are valuable mathematical problems. Terence Tao's work in signal processing made LTE and cellular communications possible because it used less bandwidth. That's worth billions. How much of "pacing" is vertical integration and solving valuable problems versus safety? I don't believe in the safety argument. We can unplug. The bigger issue is legal liability. There are frameworks for this. You sign a contract and take responsibility. I don't believe in Terminator scenarios where agents take over the internet. It's an inefficient use of resources. AI is for dirty, tedious work. Real thinking, unfortunately, I still do myself. I haven't felt it replace that yet.
Chapter Eight: The Future of AI Labs
Mr. Bubble: AI and high compute should be used for important problems and for the people who can solve them. If a GPU can improve the probability of discovering a drug by even 10%, what is that GPU worth? The drug is worth billions. That's my main point. I'm not interested in the garbage. Let's do the real work. How much of "pacing" is about allocating resources to important problems versus regulatory capture?
Logan Jastremski: I think it's mostly SCOP. The models are good but not super intelligent. I don't think we've reached that point.
Mr. Bubble: If I have a certain amount of compute, running things expensively could serve more users at $20/month or discover new materials. I'm not against the concept of pacing, just the narrative. We're bottlenecked by human effort. We have intuition but must verify. Let's use AI for important problems. Regarding open vs. closed source, running open-source models takes technical sophistication, which is why it's a big business. Open source commoditizes programming. GPT-6 is good at video games. But the frontier is pushing beyond computers—discovering new things that matter more than apps. AI is making us better, not just itself. If "pacing the frontier" means fewer users and more focus on valuable problems, it changes the business model fundamentally. But we don't know. If I could invest at a trillion-dollar valuation in an AI lab, would I? I believe in owning the hardware layer, the railroad. Anthropic and OpenAI will compete to the death. I'd rather own the companies with 10-20 year moats in core hardware. Even owning GPUs and being great at running data centers is valuable.
Chapter Nine: Conclusion and Where Value Flows Next
Mr. Bubble: Most investors will binge on these names. A good rule of thumb: if a company disappeared for six months, how bad would it be? TSMC would be terrible. NVIDIA would be terrible. Sandisk we might manage. The best companies give you no reason to leave. If an AI lab starts doing that, it'll be interesting, but they're not that smart yet. I'm looking forward to their IPOs. They'll be great short targets and great portfolio beta hedges.
Logan Jastremski: So you continue to be long hardware. It's really about hedging and adding new gigawatts. The "pacing the frontier" narrative seems like a farce because they can't pause the data centers they're building.
Mr. Bubble: It's "pacing the frontier" for us, not for themselves.
Logan Jastremski: Exactly. They talk about adding more gigawatts. Building data centers is a multi-year commitment. If they were serious, they'd cancel the capex. That's not happening. I'm very bullish. We have rate issues and Iran, but with AI we'll discover drugs, aggregate the right researchers, and find more efficient tech. We'll have a fast takeoff. Five to ten drugs will justify all the capex.
Logan Jastremski: Thank you for coming on the show.
Mr. Bubble: Thank you for having me.