Essays
  • data-ai
  • history

Everybody wants to rule the world

The mathematics of thinking machines was finished decades before anyone could afford to run it. A trace from Galvani's twitching frog to the graphics card — and the question the whole enterprise set aside in order to begin.

Why did the machines learn to talk in 2023 rather than in 1975? Why did a field that spent four decades as a cautionary tale become, more or less overnight, the most consequential technology of our working lives? What, precisely, was discovered in the interval?

The honest answer is less flattering to our moment than we might wish. Almost nothing in the present moment is a new idea. The mathematics that describes information was finished in 1948; the mathematics that describes a learning machine was finished by 1958; the algorithm that trains a deep one was in print by 1986. What arrived later was not insight. What arrived was the ability to afford the insight we already had. The ideas were finished. They were waiting on the hardware.

That is the claim, anyway, and it seems to me worth tracing carefully, because the lineage runs further back than most tellings admit — back past the transistor, past the telephone, to a dissected frog on a table in Bologna.

There is a second story tangled up in that one, and it is the story I find harder to shake. The people at the beginning of this lineage were trying to understand something: what a nerve is, what information is, what thought might turn out to be. The people who eventually paid for the machine were trying to win something — a game first, and then a fortune. Both stories arrive in the same room at more or less the same hour, and the room is the one we are standing in now.

The twitching frog

In the early 1780s, in a laboratory cluttered with the apparatus of a gentleman scientist, a dead frog’s leg kicked. Luigi Galvani had been working near a machine that threw off sparks, and he noticed — the way a careful person notices — that the muscle moved when it had no business moving. He pursued the observation for a decade, and by 1791 he had a theory: animals carry an electricity of their own, a vital fluid resident in the tissue, distinct from the electricity of storms and spark-machines. He called it animal electricity.

Alessandro Volta thought this was nonsense, and said so at length. The electricity, he argued, came from the two dissimilar metals that Galvani had used to touch the wet tissue; the frog was not a source at all, merely a conductor with rather better publicity than it deserved. To prove the point, Volta stacked discs of zinc and silver with brine-soaked cloth between them and produced a steady current from an apparatus containing no animal whatsoever. He had built the first battery, in 1800, essentially as a rebuttal.

They were each about half right, which is the most any of us can hope for. Volta was correct that his pile needed no frog; Galvani was correct that the frog was doing something electrical on its own account. The dispute is the part worth keeping. We tend to narrate discovery as a moment of arrival — a leg twitches, a truth is revealed — when the actual mechanism is closer to two people being usefully wrong at each other for twenty years. Volta chased his half of the argument into a device that would power the next century. Galvani’s half took another hundred and fifty years to come good.

The membrane, and what crosses it

What Galvani could not have known is that the frog’s electricity is not electricity in Volta’s sense at all.

The first real proposal came from Julius Bernstein in 1902: the signal in a nerve is not a fluid running down a pipe, nor a current running down a wire, but a difference in concentration of charged particles across a membrane. The cell spends energy — continuously, expensively — pumping sodium out and potassium in, holding itself in a state of readiness the way a drawn bow holds itself. The resting neuron is not idle; it is straining.

Confirmation waited on an animal large enough to instrument. Alan Hodgkin and Andrew Huxley found one in the squid, whose giant axon — the fiber that fires its escape reflex — runs nearly a millimeter across, wide enough to thread an electrode into. Their papers of 1952 gave the field a set of differential equations that still hold up, and the Nobel followed in 1963.

The picture they produced is worth sitting with, because it reframes what a neuron is. A wire carries a signal by conducting it. A neuron carries a signal by letting go — gates in the membrane snap open, the gradient it has been paying to maintain collapses locally, and the collapse propagates, triggering the next patch of membrane to release in turn. The signal is not the charge. The signal is the failure, in sequence, of a barrier the cell has been holding shut. What travels down the axon is closer to a rumor than to a current.

By mid-century, then, the nerve impulse had become a quantity: measurable, modelable, written down in equations. That turns out to matter enormously, because a thing that can be written down can be imitated by something that is not a nerve.

What a bit is

The second thread begins in a telephone company, and it begins with a practical challenge: how much can we push down a wire, and how would we know?

Harry Nyquist took up the question at Bell Telephone Laboratories in 1924, asking what limits the speed of a telegraph line. Ralph Hartley, at the same institution in 1928, pressed further and reached for a definition of information itself — a logarithmic measure, counting not what a message meant but how many messages it might have been. The pieces were on the table for twenty years.

Claude Shannon assembled them in 1948, in a paper called A Mathematical Theory of Communication, and the assembly was so complete that the field arrived fully grown. His central move was to sever information from meaning altogether. The quantity of information in a message has nothing to do with its importance, its truth, or its subject; it has to do with how thoroughly it resolves our uncertainty about which message we were going to get. The unit — the bit, a name Shannon credits to his colleague John Tukey — is not a unit of stuff. It is a unit of surprise. One bit is the amount of surprise in a fair coin landing.

Here we encounter one of the strangest coincidences in the history of science. The formula Shannon derived for the average surprise of a message source is, up to a constant, the formula Ludwig Boltzmann had derived in the 1870s for the entropy of a gas. Two fields, working from opposite ends of the universe — one counting the ways molecules can be arranged in a box, the other counting the ways a sentence can arrive down a wire — wrote down the same equation. Both, on inspection, were counting the same thing: ignorance. The number of arrangements we cannot distinguish between.

There is a widely repeated story that John von Neumann advised Shannon to call his quantity entropy on the grounds that nobody understands entropy, so he would win any argument about it. The story is secondhand and probably too good to be true; I include it knowing full well it may be a fable, but it’s a delightful “bit” of apocrypha.

The neuron as arithmetic

The two threads meet in 1943, and they meet in an unlikely pair of hands.

Warren McCulloch was a neurophysiologist with a philosopher’s temperament. Walter Pitts was, by the time they met, a homeless teenager; the story goes that he had read the Principia Mathematica in a library at twelve and written to Bertrand Russell about errors he had found in it. Their paper — A Logical Calculus of the Ideas Immanent in Nervous Activity — proposed that a neuron could be treated as a simple arithmetic device: it sums its inputs, compares the sum against a threshold, and fires or does not fire. Nothing about ions, nothing about membranes, nothing about squid. An abstraction, deliberately impoverished.

From that impoverished unit they proved something remarkable. A network of such units, suitably arranged, can compute any proposition of logic whatsoever. Thought, or at least the formal skeleton of it, was in principle available to an assembly of switches.

Donald Hebb supplied the missing half in 1949 with a proposal about learning: when one cell repeatedly participates in firing another, the connection between them strengthens. Cells that fire together, wire together. Learning, on this account, is not the acquisition of facts but the adjustment of connection strengths — a claim that reduces an enormous philosophical problem to a numerical one.

Frank Rosenblatt put the pieces together at the Cornell Aeronautical Laboratory in 1957 and 1958. His perceptron had what McCulloch and Pitts had lacked — a rule for adjusting the weights in response to error — and, better still, it had a body. The Mark I Perceptron was a physical machine: a grid of four hundred photocells for an eye, and weights stored as the settings of potentiometers, which motors physically turned as the machine learned. One could stand in the room and watch it change its mind.

The press response of 1958 will feel familiar to anyone reading a technology section this morning. The New York Times reported the Navy’s expectation of a machine that would walk, talk, see, write, reproduce itself, and be conscious of its existence. The hype cycle is not a modern invention; it is at least as old as the thing it hypes.

A correction came in 1969, when Marvin Minsky and Seymour Papert published Perceptrons and proved, rigorously, that a single layer of these units cannot compute even the exclusive-or — the humble logical function that answers one or the other, not both. Funding evaporated; a winter set in that lasted the better part of two decades.

The detail that matters, and that most retellings skip, is that the proof was far narrower than its reception. Minsky and Papert had shown a limit on single-layer networks. Stacking layers was already understood to dissolve the limitation. The obstacle was never that multi-layer networks could not work; the obstacle was that nobody knew how to construct and train one.

Then, in 1986, a modest proposal appeared. Backpropagation — anticipated by Paul Werbos in 1974 and brought to the field’s attention by David Rumelhart, Geoffrey Hinton, and Ronald Williams — supplied a method for assigning “blame” backward through the layers, so that every parameter weight in a deep network could learn from the network’s error at the output. The mathematics amounts to little more than the chain rule applied with patience, which is precisely what makes it clever rather than merely complicated; the potential sitting inside it was enormous, and took another quarter century to become obvious to anyone.

If this is the right read, 1986 is the year the mystery ran out. The unit was specified in 1943, the learning principle in 1949, the trainable machine in 1958, the deep training algorithm in 1986. What followed was not a conceptual gap. What followed was twenty-six years of not having enough arithmetic.

The database that had to wait

Before we blame the silicon, a control case is worth examining, from a corner of computing where nobody was chasing intelligence at all.

Edgar Codd published A Relational Model of Data for Large Shared Data Banks in 1970 and, in about fifteen pages, gave the industry a complete mathematical account of what it means to ask a question of stored data. His relational algebra was closed, general, and provably sufficient; it also proposed something close to heresy, which was that a programmer ought to describe what they wanted rather than how to go and fetch it.

The objection was never that Codd was wrong. The objection, sustained for most of a decade by the people running the incumbent navigational systems, was that he was slow. Handwritten retrieval code beat a general query engine, and everyone knew it.

What closed the gap was a decade of unglamorous engineering — System R at IBM, Ingres at Berkeley, Oracle’s commercial release in 1979 — and, at the heart of it, the cost-based query optimizer: a program that reads a declarative request and works out a good way to execute it. The formalism had been waiting on a compiler good enough to make it cheap, and on machines whose cycles had become plentiful enough that a little inefficiency stopped mattering.

The tempting flourish here is that Codd’s paper preceded the Intel 4004 by a year, and that the microprocessor rescued the relational model. It is too neat, and it is not true; no relational database ever ran on a four-bit calculator chip, and System R ran on mainframes. What the 4004 announced was not a platform but a trajectory — the beginning of a collapse in the price of computation that would eventually make Codd’s elegance affordable. The idea was correct on arrival and unaffordable on arrival, and only one of those two conditions changed.

Teenagers, and the accident of the polygon

All of which brings us, by a route nobody planned, to first-person shooters.

When id Software shipped Wolfenstein 3D in 1992, Doom the following year, and then Quake in 1996, the demand they created had nothing to do with AI research. What the market wanted was a corridor to run down at thirty-five frames a second, and something waiting at the end of it — Nazis first, then demons, each escalation quietly amounting to a hardware requirement — rendered convincingly enough to be genuinely alarming.

The computational shape of that problem is peculiar and, as it turns out, somewhat lucky. Rendering a scene means performing the same small arithmetic operation on an enormous number of independent items — this vertex, that pixel, none of them needing to know what the others are doing. Problems of this kind have a name in the trade: embarrassingly parallel, meaning the work divides so cleanly that splitting it up requires no cleverness at all. Such problems are miserably served by a fast general-purpose processor, or CPU, doing one thing at a time, and beautifully served by a thousand feeble processors doing a thousand things at once.

Nobody would have funded that second machine on the merits. It is useless for spreadsheets, useless for databases, useless for nearly everything a business does. What funded it was a consumer market of millions of teenagers, each willing to pay a couple of hundred dollars for a better monster to fight: 3dfx’s Voodoo arrived in 1996, NVIDIA’s GeForce 256 in 1999 — the novel technology had the term GPU attached — and, in 2001, programmable shaders, which let a developer write arbitrary small programs to run on all those parallel units. The fixed-function rendering pipeline had quietly become a general-purpose parallel computer, sold at consumer volumes and consumer prices.

The academics noticed. CUDA gave the accidental supercomputer a proper programming interface in 2007; Rajat Raina, Anand Madhavan, and Andrew Ng demonstrated GPU-trained deep networks in 2009; and in 2012, Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton trained AlexNet on two consumer gaming cards and won an image recognition contest by a margin that ended the argument.

Rosenblatt’s motors, turning potentiometers one at a time, had become several thousand arithmetic units turning weights at once, and the thing that paid for them was entertainment.

Speculators, and the second subsidy

Entertainment paid for the proof. Something stranger paid for the scale.

Bitcoin’s miners worked out in 2010 that the puzzle at the heart of proof-of-work — the arrangement in which computers race to solve a deliberately pointless problem, the winner earning the right to write the next page of the ledger — is embarrassingly parallel in exactly the way rendering is. The race was running on graphics cards within months, then left them again by 2013 for application-specific chips built to do that one puzzle and nothing else.

One interesting turn came with Ethereum in 2015, whose puzzle was purposefully designed to resist such chips by demanding vast quantities of fast memory, rather than raw arithmetic. The practical effect of that design choice was to bind an entire speculative economy to ordinary consumer graphics cards for the better part of seven years.

What followed, through the manias of 2017 and 2021, is familiar to anyone who tried to buy a computer in those years. Cards sold far above list price when they could be found at all; warehouses in Sichuan and Kazakhstan filled with rack after rack of consumer hardware bought for a purpose its designers never imagined; NVIDIA eventually shipped cards deliberately hobbled at mining, hoping to shoo the speculators back toward products built for them.

We ought to be careful about what this did and did not do for the science, since the flattering version of the story is not quite the true one. The shortages hurt researchers directly — a graduate student with a modest grant was bidding against a mania, and losing. The cards themselves largely did not become AI cards; the datacenter parts that train large models are a separate product line entirely. What the mania supplied was neither charity nor hardware, but proven scale.

It was a decade of enormous, sustained, indifferent demand that proved a business could be built selling parallel arithmetic to customers with no interest whatsoever in graphics — demand that financed the fabrication capacity, the high-bandwidth memory supply, and the corporate balance sheet from which the datacenter line was subsequently built. Racks of GPUs humming away at something other than pictures stopped being exotic and became an ordinary industrial fact.

The timing at the end of it verges on comedy. Ethereum abandoned proof-of-work in September of 2022, and the demand that had consumed those cards for seven years evaporated in a single afternoon. Roughly ten weeks later, ChatGPT was released to the public. I would not want to claim the one caused the other; the datacenter buildout was already well underway, and the coincidence is a coincidence.

Still, the shape of the thing is hard to ignore. Twice now, an industry with no interest in machine intelligence has been talked by profit into building precisely the machine that machine intelligence required — first by teenagers who wanted a better monster, then by speculators who wanted a better return. Capital is perfectly indifferent to what it ends up building. That indifference has been, so far, remarkably generous to us.

Both of these parallel computing threads share one thing in common: Everybody wants to rule the world.

What was actually invented

The chain, laid end to end, runs something like this. A frog’s leg twitches near a spark machine. The twitch is traced, over a century and a half, to charged particles crossing a membrane that the cell pays to hold shut. The membrane is abstracted into arithmetic — sum the inputs, compare to a threshold, fire. Information is severed from meaning and given a unit of its own, which turns out to be the same quantity physicists had been using to count the arrangements of a gas. The arithmetic units are stacked into layers, given a rule for learning from their own error, and then set aside for a generation, because running them at any useful size would have cost more than anyone had.

A generation of teenagers, wanting a demon drawn convincingly in a 3D space, funds the exact silicon those layers required; a decade of speculators, wanting a return on a pointless puzzle, funds the industry that makes that silicon at scale. The layers are dusted off. The machines begin to talk.

Simply put, the binding constraint in this field has almost never been the idea. It has been the substrate.

I find this thought to be clarifying rather than deflating, and useful in a practical way for anyone deciding where to put money or attention. The question of the moment is rarely has anyone thought of this? Someone usually has, and often decades ago, and the paper is usually shorter and clearer than what has been written about it since.

The question that actually discriminates is what changed in the cost of doing it? The relational model waited on the optimizer and the falling price of cycles; the deep network waited on a graphics card; and in both cases the people who moved first were the ones who noticed a price change rather than the ones who had a new thought.

Which leaves an uncomfortable question I do not think I can answer, though it seems to me the right one to be asking. Somewhere in the literature of the last fifty years, there is very likely a formalism that is correct, complete, well-understood, and dismissed as hopelessly impractical — waiting, as Rosenblatt’s layers waited, for a machine that has not been built yet, in service of a market that has nothing to do with it.

We are not smarter than Rosenblatt. We are better supplied.

The question we set aside

There is something in the motives of this story that deserves more attention than it usually receives.

Everyone in the first half was trying to understand. Galvani wanted to know why a dead leg moves. Bernstein, Hodgkin, and Huxley wanted to know what a nerve actually does. Shannon wanted to know what information is. McCulloch and Pitts wanted to know what thought is, and were willing to insult the biology to find out. The whole enterprise was world-modeling in the oldest sense: write down what we believe we know, build the thing the writing describes, and discover whether the writing holds.

There is a property of that practice we tend to underrate. A model built to explain can also be run. Understanding, written down precisely enough, becomes machinery whether or not anybody set out to build a machine.

The second half of the story runs on entirely different fuel. Nobody buying a Voodoo card in 1996 was trying to understand thought; they wanted to win a game. Nobody filling a warehouse in Kazakhstan in 2021 was trying to understand thought; they wanted to win money. Curiosity specified the machine, and appetite built it — first the appetite to rule a rendered world, then the appetite to rule this one.

Arthur C. Clarke and Stanley Kubrick put HAL 9000 on screen in 1968, one year before Minsky and Papert proved the field into its winter. The culture had imagined the conscious machine before the mathematics was declared a dead end, which strikes me as a rather tender ordering of events. We knew what we wanted long before we knew whether we could have it, and we were wrong about the timing in both directions at once.

Which returns us to 1948. Shannon’s seminal move — the one that made all of this possible — was to sever information from meaning, heralding the arrival of a truly postmodern philosophy. A binary digit does not care what it says. That severance is what allowed information to become a quantity; the quantity is what allowed thought to become arithmetic; and arithmetic, in the end, is what a graphics card is for.

The machine we built on that foundation is now the thing pressing us hardest on the question the foundation set aside. What does it mean? Is there anything it is like to be this? I have no answer, and I confess some suspicion of anyone who reports one with confidence, since the tools that carried us here were built by declining the question on purpose.

Perhaps that is the uncomfortable shape of human endeavor. We set meaning aside in order to make progress; we chased the progress with entertainment, and then with money; and we have arrived, sooner than expected and by a route nobody planned, back at the question we had agreed not to ask.

What echoes in my own head, at least, to a singable tune, no less, is that everybody wants to rule the world. In terms of AGI, if we’re truly beginning to hit Kurzweil’s singularity, the forboding question is: will we have company?