This piece is an expansion of a talk I gave in 2023. The framework comes up often enough that it seemed worth recording. It leans pretty heavily on metaphor, but until metascience develops better models, we need to lean on metaphors to describe important principles of research organizations and ecosystems.
Here’s a fairly uncontroversial statement: Healthy research ecosystems are good.
The problem is that, at all scales — from individual labs to countries — what it means for a research ecosystem to be “healthy” is incredibly hard to pin down. A healthy research ecosystem is like porn: you only know them when you see it. (And often only in retrospect!)
Beyond “smart, serious people doing good work” there don’t seem to be many similarities between, say, Bell Labs, the Manhattan Project, and the Boyden Lab at MIT — all of which are/were clearly healthy research ecosystems.
However, staring at an eye-crossingly large number of examples of both healthy and unhealthy research ecosystems has left a sort of afterimage. It consists of three components that I describe metaphorically as:
“Bubbling cauldrons” of a critical density of small-scale, bottom-up, practitioner-directed work.
Transient, top-down-directed “strike teams.”
Mechanisms to keep the “cauldrons” at a boiling temperature, nucleate bubbles that can become strike teams, recognize when one of those bubbles is ripe to become a strike team, and then rapidly form/scale a team around it.
This pattern repeats on multiple scales: at a national or regional scale with multiple organizations fulfilling each role, within a large organization like Bell Labs or DeepMind with teams filling roles, or even within single academic labs.
Putting all that together: healthy research ecosystems are a fractal combination of bubbling cauldrons and transient strike teams. The rest of this piece unpacks each chunk of that mouthful.
Bubbling cauldrons
A quick science aside that will be relevant to the extended analogy: Liquids need two elements to reach a roiling boil: sufficient heat energy and nucleation sites for bubbles to form.
Healthy research ecosystems seem to always have a mass of people doing relatively small projects that they think they should be doing. These researchers communicate with one another and come to some consensus on what the important problems are, but there’s no entity coordinating their work. There’s a vibrancy and energy to the community.
It can take a long time for a particular community or domain to come to a boil. Modern physics had been “heating up” for several decades before it came to a boil in the 1930s or so. Even with enough energy, an ecosystem needs “nucleation points”: physical gathering places, leaders/connectors, and systems like shared epistemologies that enable people to build on each other’s work. If everybody is going off in their own direction, even tons of energy won’t lead to anything good.
It’s also hard to control the factors that go into great bubbling cauldrons from the top down beyond just giving people the tools to “do stuff.” People have tried to create a million “ecosystems” and failed. The hard-to-pinpoint inputs that go into creating the bubbling cauldron piece of a healthy ecosystem is, I believe, what Vannevar Bush was gesturing at when he talked about the importance of scientist-led basic science. He ignored two things:
Scientist-led was necessary but not sufficient to create the vibrant bubbling cauldrons. You also need net positive energy into the system, which decades of increasing bureaucracy and peer review have drained out.
The bubbling cauldrons were only half of the equation. As of the 1940s, companies and the government were very good at identifying bubbles and creating strike teams around them. Given that, he just assumed the strike team and mechanism pieces of healthy ecosystems were taken care of. Bluntly, most of how we organize science today is cargo-culting the mid-20th century. We have inherited the tradition of largely ignoring strike team creation, except in a few specific cases.
Transient strike teams
In addition to the bubbling cauldrons, healthy research ecosystems have mechanisms to create transient strike teams: that is, a way to temporarily coordinate a larger number of researchers than would organically get together, aimed at a single, well-defined goal. “Hey, you/we all need to be working on this goal. Yes, I realize you all have a specific set of research interests that aren’t quite this. Yes, I realize you wouldn’t ordinarily collaborate with XYZ. Yes, I realize you hate being told what to do. But here are the excellent reasons why you should temporarily subsume yourself to this larger goal.” Those reasons are some combination of glory, duty, and money.
These strike teams require great leadership. Great people act as crystallization points for supersaturated forces. “Research management” is almost anathema to bubbling cauldrons but incredibly important for strike teams. This distinction is one of the many places where this model is important for untangling ongoing arguments like whether scientists should be managed or not – the answer is “it depends on which mode they’re in!”
The transient nature of these strike teams is important. The transience lets you suspend normal rules, incentives, hierarchies, and protocols to just get the job done. A consistent thing you hear about all of <the Manhattan project, the Rad Lab in WWII, Xerox PARC> is how everybody was focused on the mission to the exclusion of all else. This state is not sustainable. Eventually people start worrying about careers, power, etc. Eventually people start violating trust or become self-serving and you need to put bureaucracies in place to prevent that over the long term.
Furthermore, you can’t get talented individuals (especially researchers) to do work that deviates from their main interests forever. Great researchers yearn to return to cauldron mode. In other words, the strike teams shouldn’t worry about organizational technical debt and hierarchies the way that perpetual organizations need to. I’ll return to more implications of strike team transience in the corollaries section below.
Fractal
Finally, healthy research ecosystems have these strike teams and bubbling cauldrons repeating at different scales: ecosystems that span multiple organizations and institutions, within an individual organization, and even within a single PI (Principal Investigator)-led lab.
At the scale of multiple organizations and institutions: 20th century physics and the Manhattan Project
The Manhattan Project is a classic example of healthy research but it didn’t exist in a vacuum. It was a large strike team built on top of the bubbling cauldron that was early 20th century physics. The Manhattan Project couldn’t have happened without self-organizing groups of a few individuals pushing forward quantum mechanics, nuclear physics, and other areas that created the scientific foundation for building the atomic bomb.
World War II was rife with research strike teams that pushed over the finish line everything from Radar to Penicillin. It’s perhaps anachronistic, but I suspect that the success of these strike teams and the realization that they were built on bubbling cauldrons of scientist-initiated work is what informed The Endless Frontier Framework/The linear model of innovation that dominates today’s research ecosystem. It is suitcase-handle word-ing/cargo-cult-ing how good strike team/cauldron dynamics work. That is, we have encoded structures that assumed components of a healthy ecosystem that have since withered.
Today, you can also think of DARPA as a strike team generator. It can’t exist without the bubbling cauldrons of research at universities and other research organizations.
At the scale of large research organizations: Bell Labs
Notoriously effective research organizations like Bell Labs also worked via strike teams and bubbling cauldrons. The majority of the research division was organized into relatively autonomous groups of less than five people. Occasionally, leadership would identify high priority projects – like transistors after the initial point-contact proof of principle – and pull groups of dozens of researchers together to work on a single output.
(A quick aside on Bell Labs’ Research Division) This is a reminder that when most people talk about “Bell Labs” they are referring just to the research area, which was actually just a small piece of a larger organization. The research division created things like the transistor and discovered the cosmic microwave background. It was also 6% of the entire org. Most people working at Bell Labs were doing much more prosaic things like burying telephone poles for a decade to see which coatings prevented rot the best.
(End Aside)
In its heyday, DeepMind followed a similar pattern. Most researchers are in small, self-organizing groups, but projects like AlphaGo and later AlphaFold pulled in dozens of people based on the direction of Demis Hassabis, DeepMind’s CEO. (I do not know whether this is still the case in 2026 after the merger with Google Brain).
At the scale of individual labs: The Baker or Church Lab
Perhaps it’s a bit of a stretch but you can map the strike team - cauldron model onto many technical academic labs above a certain size – say the Baker Lab or Church Lab – but I believe this applies to basically any successful academic lab. These labs often have many graduate students working on individual or small-group projects with loose oversight from the PI, but occasionally the PI pulls a lot of grad students onto a single project (perhaps because it seems very promising or because of a large grant).
Corollaries
So what? The only way this cute model matters is if it bottoms out in ways of diagnosing and improving research ecosystems. Here are some actionable corollaries that (assuming this model is accurate) should lead to better research. Of course, many of these are like dieting and exercise for losing weight – easy to say, hard to do.
One corollary is that your levers for improving an ecosystem fall into two buckets:
Driving more heat into the cauldron. Heating the cauldron looks like moves that help researchers explore their own ideas, run down hunches, and spontaneously collaborate. Things like grants with broad calls, researcher-led workshops, etc.
Building better mechanisms for strike team creation or creating strike teams directly. Enabling more strike teams looks like empowering and supporting opinionated program leaders to work on big things, reducing the friction to starting ambitious programs, and forcing collaboration between relevant players in the ecosystem towards specific goals.
The trick is knowing which intervention an ecosystem needs. People often try to summon strike teams from a cold cauldron -- this is the failure mode of places that try to be the next technology hub via startups (which, in the beginning, are a form of strike team) -- or they try to build ambitious programs in research disciplines where there really aren’t any great initial results. Similarly, it can feel good to add additional heat to an already critically hot cauldron by giving out small grants or other small increments of support when what they really need is large chunks of support around opinionated leaders.
Watched pots never boil. It’s not strictly true for real or metaphorical pots, but is directionally right -- cauldrons coming to a bubble can be encouraged but not rushed. Because it is a bottom-up process, interventions can’t speed it up on a deterministic timeline, which is antithetical to the way many institutions think about impact.
Strike team ephemerality means that strike teams can metastasize if they outlive their usefulness. That, of course, raises the question “how long should strike teams last?” The answer will depend heavily on the actual goal of the team.
In order to support strike team ephemerality, it’s important to make it easy for the people on a strike team to transition out of it and for its supporters to call it a win. Otherwise, there will be strong incentives on all sides towards scope creep and trying to become a permanent thing. Operation Warp Speed is a great example of doing this well.
Strike teams need both the autonomy to run headlong at their goals without constant outside intervention but also need the authority to, bluntly, tell people what to do. Insufficient autonomy and authority lead to several failure modes:
Needing to check off a bunch of boxes that are unrelated to their goals -- this often happens through government bureaucracy. (Autonomy failure)
Uncertainty over whether they will actually be supported through the extent of their mission which leads to wasting a lot of time on shorter term demos and reports to justify their existence or try to raise additional funding. (Autonomy failure)
Becoming little more than an excuse/funding source for a bunch of individuals or small groups to do whatever they think is most interesting. This is the failure mode of many university centers and institutes which are basically just a bunch of normal PI-led labs in a trench coat. (Authority failure)
Strike teams need to be created quickly. There is often a short time window where the right people and resources align with enough momentum to start a big hard thing. Many strike-team-creation-mechanisms fail because the process is drawn out -- the best people don’t just sit around ready to drop everything for half a year until a review process is completed; they will go for a different all-consuming opportunity.
Boiling requires a lot of stochastic motion. That is, you need a lot of people working on a lot of uncorrelated things in a general area, many of which you should expect to fail. This corollary is more subtle than just banging the tired drum of supporting “high risk” projects -- it means putting less pressure on boiling-cauldron-work to justify itself and supporting things that are more weird or off-the-beaten path with the expectation that most of them will go nowhere.
Camelot!
I’ll close with a meta-observation.
Obviously, bubbling cauldrons and strike teams is only a model. But models can be powerful tools for exploring systems and ideas that are “off the observable map.” Models are what let you predict the outcomes of actions that have never been taken before. Of course, all models are wrong, but some of them are useful.
The fuzzily-defined field of metascience has shockingly few models. We’re not even in the “all heavenly objects are embedded in spheres that revolve around the earth” era, let alone heliocentrism. What models we do have are generally either implicit “folk models” or structures that were created for bureaucratic legibility rather than understanding. We can change that!



