本周精选(第 313 周) | Azimuth --- This Week’s Finds (Week 313) | Azimuth
Here’s the third and final part of my interview with Eliezer Yudkowsky. We’ll talk about three big questions… roughly these:
- How do you get people to work on potentially risky projects in a safe way?
- Do we understand ethics well enough to build “Friendly artificial intelligence”?
- What’s better to work on, artificial intelligence or environmental issues?
So, with no further ado:
JB: There are decent Wikipedia articles on “optimism bias” and “positive illusions”, which suggest that unrealistically optimistic people are more energetic, while more realistic estimates of success go hand-in-hand with mild depression. If this is true, I can easily imagine that most people working on challenging projects like quantum gravity (me, 10 years ago) or artificial intelligence (you) are unrealistically optimistic about our chances of success.
Indeed, I can easily imagine that the first researchers to create a truly powerful artificial intelligence will be people who underestimate its potential dangers. It’s an interesting irony, isn’t it? If most people who are naturally cautious avoid a certain potentially dangerous line of research, the people who pursue that line of research are likely to be less cautious than average.
I’m a bit worried about this when it comes to “geoengineering”, for example—attempts to tackle global warming by large engineering projects. We have people who say “oh no, that’s too dangerous”, and turn their attention to approaches they consider less risky, but that may leave the field to people who underestimate the risks.
So I’m very glad you are thinking hard about how to avoid the potential dangers of artificial intelligence—and even trying to make this problem sound exciting, to attract ambitious and energetic young people to work on it. Is that part of your explicit goal? To make caution and rationality sound sexy?
EY: The really hard part of the problem isn’t getting a few smart people to work on cautious, rational AI. It’s admittedly a harder problem than it should be, because there’s a whole system out there which is set up to funnel smart young people into all sorts of other things besides cautious rational long-term basic AI research. But it isn’t the really hard part of the problem.
The scary thing about AI is that I would guess that the first AI to go over some critical threshold of self-improvement takes all the marbles—first mover advantage, winner take all. The first pile of uranium to have an effective neutron multiplication factor greater than 1, or maybe the first AI smart enough to absorb all the poorly defended processing power on the Internet—there’s actually a number of different thresholds that could provide a critical first-mover advantage.
And it is always going to be fundamentally easier in some sense to go straight all out for AI and not worry about clean designs or stable self-modification or the problem where a near-miss on the value system destroys almost all of the actual value from our perspective. (E.g., imagine aliens who shared every single term in the human utility function but lacked our notion of boredom. Their civilization might consist of a single peak experience repeated over and over, which would make their civilization very boring from our perspective, compared to what it might have been. That is, leaving a single aspect out of the value system can destroy almost all of the value. So there’s a very large gap in the AI problem between trying to get the value system exactly right, versus throwing something at it that sounds vaguely good.)
You want to keep as much of an advantage as possible for the cautious rational AI developers over the crowd that is just gung-ho to solve this super interesting scientific problem and go down in the eternal books of fame. Now there should in fact be some upper bound on the combination of intelligence, methodological rationality, and deep understanding of the problem which you can possess, and still walk directly into the whirling helicopter blades. The problem is that it is probably a rather high upper bound. And you are trying to outrace people who are trying to solve a fundamentally easier wrong problem. So the question is not attracting people to the field in general, but rather getting the really smart competent people to either work for a cautious project or not go into the field at all. You aren’t going to stop people from trying to develop AI. But you can hope to have as many of the really smart people as possible working on cautious projects rather than incautious ones.
So yes, making caution look sexy. But even more than that, trying to make incautious AI projects look merely stupid. Not dangerous. Dangerous is sexy. As the old proverb goes, most of the damage is done by people who wish to feel themselves important. Human psychology seems to be such that many ambitious people find it far less scary to think about destroying the world, than to think about never amounting to much of anything at all. I have met people like this. In fact all the people I have met who think they are going to win eternal fame through their AI projects have been like this. The thought of potentially destroying the world is bearable; it confirms their own importance. The thought of not being able to plow full steam ahead on their incredible amazing AI idea is not bearable; it threatens all their fantasies of wealth and fame.
Now these people of whom I speak are not top-notch minds, not in the class of the top people in mainstream AI, like say Peter Norvig (to name someone I’ve had the honor of meeting personally). And it’s possible that if and when self-improving AI starts to get real top-notch minds working on it, rather than people who were too optimistic about/attached to their amazing bright idea to be scared away by the field of skulls, then these real stars will not fall prey to the same sort of psychological trap. And then again it is also plausible to me that top-notch minds will fall prey to exactly the same trap, because I have yet to learn from reading history that great scientific geniuses are always sane.
So what I would most like to see would be uniform looks of condescending scorn directed at people who claimed their amazing bright AI idea was going to lead to self-improvement and superintelligence, but who couldn’t mount an adequate defense of how their design would have a goal system stable after a billion sequential self-modifications, or how it would get the value system exactly right instead of mostly right. In other words, making destroying the world look unprestigious and low-status, instead of leaving it to the default state of sexiness and importance-confirmingness.
JB: “Get the value system exactly right”—now this phrase touches on another issue I’ve been wanting to talk about. How do we know what it means for a value system to be exactly right? It seems people are even further from agreeing on what it means to be good than on what it means to be rational. Yet you seem to be suggesting we need to solve this problem before it’s safe to build a self-improving artificial intelligence!
When I was younger I worried a lot about the foundations of ethics. I decided that you “can’t derive an ought from an is”—do you believe that? If so, all logical arguments leading up to the conclusion that “you should do X” must involve an assumption of the form “you should do Y”… and attempts to “derive” ethics are all implicitly circular in some way. This really bothered the heck out of me: how was I supposed to know what to do? But of course I kept on doing things while I was worrying about this… and indeed, it was painfully clear that there’s no way out of making decisions: even deciding to “do nothing” or commit suicide counts as a decision.
Later I got more comfortable with the idea that making decisions about what to do needn’t paralyze me any more than making decisions about what is true. But still, it seems that the business of designing ethical beings is going to provoke huge arguments, if and when we get around to that.
Do you spend as much time thinking about these issues as you do thinking about rationality? Of course they’re linked….
EY: Well, I probably spend as much time explaining these issues as I do rationality. There are also an absolutely huge number of pitfalls that people stumble into when they try to think about, as I would put it, Friendly AI. Consider how many pitfalls people run into when they try to think about Artificial Intelligence. Next consider how many pitfalls people run into when they try to think about morality. Next consider how many pitfalls philosophers run into when they try to think about the nature of morality. Next consider how many pitfalls people run into when they try to think about hypothetical extremely powerful agents, especially extremely powerful agents that are supposed to be extremely good. Next consider how many pitfalls people run into when they try to imagine optimal worlds to live in or optimal rules to follow or optimal governments and so on.
Now imagine a subject matter which offers discussants a lovely opportunity to run into all of those pitfalls at the same time.
That’s what happens when you try to talk about Friendly Artificial Intelligence.
And it only takes one error for a chain of reasoning to end up in Outer Mongolia. So one of the great motivating factors behind all the writing I did on rationality and all the sequences I wrote on Less Wrong was to actually make it possible, via two years worth of writing and probably something like a month’s worth of reading at least, to immunize people against all the usual mistakes.
Lest I appear to dodge the question entirely, I’ll try for very quick descriptions and google keywords that professional moral philosophers might recognize.
In terms of what I would advocate programming a very powerful AI to actually do, the keywords are “mature folk morality” and “reflective equilibrium”. This means that you build a sufficiently powerful AI to do, not what people say they want, or even what people actually want, but what people would decide they wanted the AI to do, if they had all of the AI’s information, could think about for as long a subjective time as the AI, knew as much as the AI did about the real factors at work in their own psychology, and had no failures of self-control.
There’s a lot of important reasons why you would want to do exactly that and not, say, implement Asimov’s Three Laws of Robotics (a purely fictional device, and if Asimov had depicted them as working well, he would have had no stories to write) or building a superpowerful AI which obeys people’s commands interpreted in literal English, or creating a god whose sole prime directive is to make people maximally happy, or any of the above plus a list of six different patches which guarantee that nothing can possibly go wrong, and various other things that seem like incredibly obvious failure scenarios but which I assure you I have heard seriously advocated over and over and over again.
In a nutshell, you want to use concepts like “mature folk morality” or “reflective equilibrium” because these are as close as moral philosophy has ever gotten to defining in concrete, computable terms what you could be wrong about when you order an AI to do the wrong thing.
For an attempt at nontechnical explanation of what one might want to program an AI to do and why, the best resource I can offer is an old essay of mine which is not written so as to offer good google keywords, but holds up fairly well nonetheless:
- Eliezer Yudkowsky, Coherent extrapolated volition, May 2004.
You also raised some questions about metaethics, where metaethics asks not “Which acts are moral?” but “What is the subject matter of our talk about ‘morality’?” i.e. “What are we talking about here anyway?” In terms of Google keywords, my brand of metaethics is closest to analytic descriptivism or moral functionalism. If I were to try to put that into a very brief nutshell, it would be something like “When we talk about ‘morality’ or ‘goodness’ or ‘right’, the subject matter we’re talking about is a sort of gigantic math question hidden under the simple word ‘right’, a math question that includes all of our emotions and all of what we use to process moral arguments and all the things we might want to change about ourselves if we could see our own source code and know what we were really thinking.”
The complete Less Wrong sequence on metaethics (with many dependencies to earlier ones) is:
- Eliezer Yudkowsky, Metaethics sequence, Less Wrong, 20 June to 22 August 2008.
And one of the better quick summaries is at:
- Eliezer Yudkowsky, Inseparably right; or, joy in the merely good, Less Wrong, 9 August 2008.
And if I am wise I shall not say any more.
JB: I’ll help you be wise. There are a hundred followup questions I’m tempted to ask, but this has been a long and grueling interview, so I won’t. Instead, I’d like to raise one last big question. It’s about time scales.
Self-improving artificial intelligence seems like a real possibility to me. But when? You see, I believe we’re in the midst of a global ecological crisis—a mass extinction event, whose effects will be painfully evident by the end of the century. I want to do something about it. I can’t do much, but I want to do something. Even if we’re doomed to disaster, there are different sizes of disaster. And if we’re going through a kind of bottleneck, where some species make it through and others go extinct, even small actions now can make a difference.
I can imagine some technological optimists—singularitarians, extropians and the like—saying: “Don’t worry, things will get better. Things that seem hard now will only get easier. We’ll be able to suck carbon dioxide from the atmosphere using nanotechnology, and revive species starting from their DNA.” Or maybe even: “Don’t worry: we won’t miss those species. We’ll be having too much fun doing things we can’t even conceive of now.”
But various things make me skeptical of such optimism. One of them is the question of time scales. What if the world goes to hell before our technology saves us? What if artificial intelligence comes along toolate to make a big impact on the short-term problems I’m worrying about? In that case, maybe I should focus on short-term solutions.
Just to be clear: this isn’t some veiled attack on your priorities. I’m just trying to decide on my own. One good thing about having billions of people on the planet is that we don’t all have to do the same thing. Indeed, a multi-pronged approach is best. But for my own decisions, I want some rough guess about how long various potentially revolutionary technologies will take to come online.
What do you think about all this?
EY: I’ll try to answer the question about timescales, but first let me explain in some detail why I don’t think the decision should be dominated by that question.
If you look up “Scope Insensitivity” on Less Wrong, you’ll see that when three different groups of subjects were asked how much they would pay in increased taxes to save 2,000 / 20,000 / 200,000 birds from drowning in uncovered oil ponds, the respective average answers were $80 / $78 / $88. People asked questions like this visualize one bird, wings slicked with oil, struggling to escape, and that creates some amount of emotional affect which determines willingness to pay, and the quantity gets tossed out the window since no one can visualize 200,000 of anything. Another hypothesis to explain the data is “purchase of moral satisfaction”, which says that people give enough money to create a “warm glow” inside themselves, and the amount required might have something to do with your personal financial situation, but it has nothing to do with birds. Similarly, residents of four US states were only willing to pay 22% more to protect all 57 wilderness areas in those states than to protect one area. The result I found most horrifying was that subjects were willing to contribute more when a set amount of money was needed to save one child’s life, compared to the same amount of money saving eight lives—because, of course, focusing your attention on a single person makes the feelings stronger, less diffuse.
So while it may make sense to enjoy the warm glow of doing good deeds after we do them, we cannot possibly allow ourselves to choose between altruistic causes based on the relative amounts of warm glow they generate, because our intuitions are quantitatively insane.
And two antidotes that absolutely must be applied in choosing between altruistic causes are conscious appreciation of scope and conscious appreciation of marginal impact.
By its nature, your brain flushes right out the window the all-important distinction between saving one life and saving a million lives. You’ve got to compensate for that using conscious, verbal deliberation. The Society For Curing Rare Diseases in Cute Puppies has got great warm glow, but the fact that these diseases are rare should call a screeching halt right there—which you’re going to have to do consciously, not intuitively. Even before you realize that, contrary to the relative warm glows, it’s really hard to make a moral case for trading off human lives against cute puppies. I suppose if you could save a billion puppies using one dollar I wouldn’t scream at someone who wanted to spend the dollar on that instead of cancer research.
And similarly, if there are a hundred thousand researchers and billions of dollars annually that are already going into saving species from extinction—because it’s a prestigious and popular cause that has an easy time generating warm glow in lots of potential funders—then you have to ask about the marginal value of putting your effort there, where so many other people are already working, compared to a project that isn’t so popular.
I wouldn’t say “Don’t worry, we won’t miss those species”. But consider the future intergalactic civilizations growing out of Earth-originating intelligent life. Consider the whole history of a universe which contains this world of Earth and this present century, and also billions of years of future intergalactic civilization continuing until the universe dies, or maybe forever if we can think of some ingenious way to carry on. Next consider the interval in utility between a universe-history in which Earth-originating intelligence survived and thrived and managed to save 95% of the non-primate biological species now alive, versus a universe-history in which only 80% of those species are alive. That utility interval is not very large compared to the utility interval between a universe in which intelligent life thrived and intelligent life died out. Or the utility interval between a universe-history filled with sentient beings who experience happiness and have empathy for each other and get bored when they do the same thing too many times, versus a universe-history that grew out of various failures of Friendly AI.
(The really scary thing about universes that grow out of a loss of human value is not that they are different, but that they are, from our standpoint, boring. The human utility function says that once you’ve made a piece of art, it’s more fun to make a different piece of art next time. But that’s just us. Most random utility functions will yield instrumental strategies that spend some of their time and resources exploring for the patterns with the highest utility at the beginning of the problem, and then use the rest of their resources to implement the pattern with the highest utility, over and over and over. This sort of thing will surprise a human who expects, on some deep level, that all minds are made out of human parts, and who thinks, “Won’t the AI see that its utility function is boring?” But the AI is not a little spirit that looks over its code and decides whether to obey it; the AI is the code. If the code doesn’t say to get bored, it won’t get bored. A strategy of exploration followed by exploitation is implicit in most utility functions, but boredom is not. If your utility function does not already contain a term for boredom, then you don’t care; it’s not something that emerges as an instrumental value from most terminal values. For more on this see: “In Praise of Boredom” in the Fun Theory Sequence on Less Wrong.)
Anyway: In terms of expected utility maximization, even large probabilities of jumping the interval between a universe-history in which 95% of existing biological species survive Earth’s 21st century, versus a universe-history where 80% of species survive, are just about impossible to trade off against tiny probabilities of jumping the interval between interesting universe-histories, versus boring ones where intelligent life goes extinct, or the wrong sort of AI self-improves.
I honestly don’t see how a rationalist can avoid this conclusion: At this absolutely critical hinge in the history of the universe—Earth in the 21st century—rational altruists should devote their marginal attentions to risks that threaten to terminate intelligent life or permanently destroy a part of its potential. Those problems, which Nick Bostrom named “existential risks“, have got all the scope. And when it comes to marginal impact, there are major risks outstanding that practically no one is working on. Once you get the stakes on a gut level it’s hard to see how doing anything else could be sane.
So how do you go about protecting the future of intelligent life? Environmentalism? After all, there are environmental catastrophes that could knock over our civilization… but then if you want to put the whole universe at stake, it’s not enough for one civilization to topple, you have to argue that our civilization is above average in its chances of building a positive galactic future compared to whatever civilization would rise again a century or two later. Maybe if there were ten people working on environmentalism and millions of people working on Friendly AI, I could see sending the next marginal dollar to environmentalism. But with millions of people working on environmentalism, and major existential risks that are completely ignored… if you add a marginal resource that can, rarely, be steered by expected utilities instead of warm glows, devoting that resource to environmentalism does not make sense.
Similarly with other short-term problems. Unless they’re little-known and unpopular problems, the marginal impact is not going to make sense, because millions of other people will already be working on them. And even if you argue that some short-term problem leverages existential risk, it’s not going to be perfect leverage and some quantitative discount will apply, probably a large one. I would be suspicious that the decision to work on a short-term problem was driven by warm glow, status drives, or simple conventionalism.
With that said, there’s also such a thing as comparative advantage—the old puzzle of the lawyer who works an hour in the soup clinic instead of working an extra hour as a lawyer and donating the money. Personally I’d say you can work an hour in the soup clinic to keep yourself going if you like, but you should also be working extra lawyer-hours and donating the money to the soup clinic, or better yet, to something with more scope. (See “Purchase Fuzzies and Utilons Separately” on Less Wrong.) Most people can’t work effectively on Artificial Intelligence (some would question if anyone can, but at the very least it’s not an easy problem). But there’s a variety of existential risks to choose from, plus a general background job of spreading sufficiently high-grade rationality and existential risk awareness. One really should look over those before going into something short-term and conventional. Unless your master plan is just to work the extra hours and donate them to the cause with the highest marginal expected utility per dollar, which is perfectly respectable.
Where should you go in life? I don’t know exactly, but I think I’ll go ahead and say “not environmentalism”. There’s just no way that the product of scope, marginal impact, and John Baez’s comparative advantage is going to end up being maximal at that point.
Which brings me to AI timescales.
If I knew exactly how to make a Friendly AI, and I knew exactly how many people I had available to do it, I still couldn’t tell you how long it would take because of Product Management Chaos.
As it stands, this is a basic research problem—which will always feel very hard, because we don’t understand it, and that means when our brain checks for solutions, we don’t see any solutions available. But this ignorance is not to be confused with the positive knowledge that the problem will take a long time to solve once we know how to solve it. It could be that some fundamental breakthrough will dissolve our confusion and then things will look relatively easy. Or it could be that some fundamental breakthrough will be followed by the realization that, now that we know what to do, it’s going to take at
least another 20 years to do it.
I seriously have no idea when AI is going to show up, although I’d be genuinely and deeply shocked if it took another century (barring a collapse of civilization in the meanwhile).
If you were to tell me that as a Bayesian I have to put probability distributions on things on pain of having my behavior be inconsistent and inefficient, well, I would actually suspect that my behavior is inconsistent. But if you were to try and induce from my behavior a median expected time where I spend half my effort planning for less and half my effort planning for more, it would probably look something like 2030.
But that doesn’t really matter to my decisions. Among all existential risks I know about, Friendly AI has the single largest absolute scope—it affects everything, and the problem must be solved at some point for worthwhile intelligence to thrive. It also has the largest product of scope of marginal impact, because practically no one is working on it, even compared to other existential risks. And my abilities seem applicable to it. So I may not like my uncertainty about timescales, but my decisions are not unstable with respect to that uncertainty.
JB: Ably argued! If I think of an interesting reply, I’ll put it in the blog discussion. Thanks for your time.
The best way to predict the future is to invent it. – Alan Kay
My friend John Barrett pointed out that the Environmental and Sustainability Institute at Exeter has jobs for mathematicians and statisticians who “combine research expertise in areas such as computational statistics, data modelling, system dynamics, control, optimization and/or computation with vision and innovation so as to transform research across the environment and sustainability agenda”.
I got his email today; the deadline has already passed, but maybe you can slip under the wire:
This entry was posted on Wednesday, March 30th, 2011 at 1:09 am and is filed under jobs. You can follow any responses to this entry through the RSS 2.0 feed. You can leave a response, or trackback from your own site.
Today I’d like to start telling you about some research Jacob Biamonte and I are doing on stochastic Petri nets, quantum field theory, and category theory. It’ll take a few blog entries to cover this story, which is part of a larger story about network theory.
Stochastic Petri nets are one of many different diagrammatic languages people have evolved to study complex systems. We’ll see how they’re used in chemistry, molecular biology, population biology and queuing theory, which is roughly the science of waiting in line. If you’re a faithful reader of this blog, you’ve already seen an example of a Petri net taken from chemistry:

It shows some chemicals and some reactions involving these chemicals. To make it into a stochastic Petri net, we’d just label each reaction by a nonnegative real number: the reaction rate constant, or rate constant for short.
In general, a Petri net will have a set of states, which we’ll draw as yellow circles, and a set of transitions, which we’ll draw as blue rectangles. Here’s a Petri net from population biology:
Now, instead of different chemicals, the states are different species. And instead of chemical reactions, the transitions are processes involving our species.
This Petri net has two states: rabbit and wolf. It has three transitions:
- In birth, one rabbit comes in and two go out. This is a caricature of reality: these bunnies reproduce asexually, splitting in two like amoebas.
- In predation, one wolf and one rabbit come in and two wolves go out. This is a caricature of how predators need to eat prey to reproduce.
- In death, one wolf comes in and nothing goes out. Note that we’re pretending rabbits don’t die unless they’re eaten by wolves.
If we labelled each transition with a rate constant, we’d have a stochastic Petri net.
To make this Petri net more realistic, we’d have to make it more complicated. I’m trying to explain general ideas here, not realistic models of specific situations. Nonetheless, this Petri net already leads to an interesting model of population dynamics: a special case of the so-called ‘Lotka-Volterra predator-prey model’. We’ll see the details soon.
More to the point, this Petri net illustrates some possibilities that our previous example neglected. Every transition has some ‘input’ states and some ‘output’ states. But a state can show up more than once as the output (or input) of some transition. And as we see in ‘death’, we can have a transition with no outputs (or inputs) at all.
But let me stop beating around the bush, and give you the formal definitions. They’re simple enough:
Definition. A Petri net consists of a set of states and a set of transitions, together with a function
saying how many copies of each state shows up as input for each transition, and a function
saying how many times it shows up as output.
Definition. A stochastic Petri net is a Petri net together with a function
giving a rate constant for each transition.
Starting from any stochastic Petri net, we can get two things. First:
- The master equation. This says how the probability that we have a given number of things in each state changes with time.
Since stochastic means ‘random’, the master equation is what gives stochastic Petri nets their name. The master equation is the main thing I’ll be talking about in future blog entries. But not right away!
Why not?
In chemistry, we typically have a huge number of things in each state. For example, a gram of water contains about water molecules, and a smaller but still enormous number of hydroxide ions (OH–), hydronium ions (H3O+), and other scarier things. These things blunder around randomly, bump into each other, and sometimes react and turn into other things. There’s a stochastic Petri net describing all this, as we’ll eventually see. But in this situation, we don’t usually want to know the probability that there are, say, exactly hydronium ions. That would be too much information! We’d be quite happy knowing the expected value of the number of hydronium ions, so we’d be delighted to have a differential equation that says how this changes with time.
And luckily, such an equation exists—and it’s much simpler than the master equation. So, today we’ll talk about:
- The rate equation. This says how the expected number of things in each state changes with time.
But first, I hope you get the overall idea. The master equation is stochastic: at each time the number of things in each state is a random variable taking values in , the set of natural numbers. The rate equation is deterministic: at each time the expected number of things in each state is a non-random variable taking values in , the set of nonnegative real numbers. If the master equation is the true story, the rate equation is only approximately true—but the approximation becomes good in some limit where the expected value of the number of things in each state is large, and the standard deviation is comparatively small.
If you’ve studied physics, this should remind you of other things. The master equation should remind you of the quantum harmonic oscillator, where energy levels are discrete, and probabilities are involved. The rate equation should remind you of the classical harmonic oscillator, where energy levels are continuous, and everything is deterministic.
When we get to the ‘original research’ part of our story, we’ll see this analogy is fairly precise! We’ll take a bunch of ideas from quantum mechanics and quantum field theory, and tweak them a bit, and show how we can use them to describe the master equation for a stochastic Petri net.
Indeed, the random processes that the master equation describes can be drawn as pictures:

This looks like a Feynman diagram, with animals instead of particles! It’s pretty funny, but the resemblance is no joke: the math will back it up.
I’m dying to explain all the details. But just as classical field theory is easier than quantum field theory, the rate equation is simpler than the master equation. So we should start there.
The rate equation
If you hand me a stochastic Petri net, I can write down its rate equation. Instead of telling you the general rule, which sounds rather complicated at first, let me do an example. Take the Petri net we were just looking at:

We can make it into a stochastic Petri net by choosing a number for each transition:
- the birth rate constant
- the predation rate constant
- the death rate constant
Let be the number of rabbits and let be the number of wolves at time . Then the rate equation looks like this:
It’s really a system of equations, but I’ll call the whole thing “the rate equation” because later we may get smart and write it as a single equation.
See how it works?
- We get a term in the equation for rabbits, because rabbits are born at a rate equal to the number of rabbits times the birth rate constant .
- We get a term in the equation for wolves, because wolves die at a rate equal to the number of wolves times the death rate constant .
- We get a term in the equation for rabbits, because rabbits die at a rate equal to the number of rabbits times the number of wolves times the predation rate constant .
- We also get a term in the equation for wolves, because wolves are born at a rate equal to the number of rabbits times the number of wolves times .
Of course I’m not claiming that this rate equation makes any sense biologically! For example, think about predation. The terms in the above equation would make sense if rabbits and wolves roamed around randomly, and whenever a wolf and a rabbit came within a certain distance, the wolf had a certain probability of eating the rabbit and giving birth to another wolf. At least it would be make sense in the limit of large numbers of rabbits and wolves, where we can treat and as varying continuously rather than discretely. That’s a reasonable approximation to make sometimes. Unfortunately, rabbits and wolves don’t roam around randomly, and a wolf doesn’t spit out a new wolf each time it eats a rabbit.
Despite that, the equations
are actually studied in population biology. As I said, they’re a special case of the Lotka-Volterra predator-prey model, which looks like this:
The point is that while these models are hideously oversimplified and thus quantitatively inaccurate, they exhibit interesting qualititative behavior that’s fairly robust. Depending on the rate constants, these equations can show either a stable equilibrium or stable periodic behavior. And we go from one regime to another, we see a kind of catastrophe called a “Hopf bifurcation”. I explained all this in week308 and week309. There I was looking at some other equations, not the Lotka-Volterra equations. But their qualitative behavior is the same!
If you want stochastic Petri nets that give quantitatively accurate models, it’s better to retreat to chemistry. Compared to animals, molecules come a lot closer to roaming around randomly and having a chance of reacting when they come within a certain distance. So in chemistry, rate equations can be used to make accurate predictions.
But I’m digressing. I should be explaining the general recipe for getting a rate equation from a stochastic Petri net! You might not be able to guess it from just one example. But I sense that you’re getting tired. So let’s stop now. Next time I’ll do more examples, and maybe even write down a general formula. But if you’re feeling ambitious, you can try this now:
Puzzle. Can you write down a stochastic Petri net whose rate equation is the Lotka-Volterra predator-prey model:
for arbitrary ? If not, for which values of these rate constants can you do it?
References
If you want to study a bit on your own, here are some great online references on stochastic Petri nets and their rate equations:
- Peter J. E. Goss and Jean Peccoud, Quantitative modeling of stochastic systems in molecular biology by using stochastic Petri nets, Proc. Natl. Acad. Sci. USA 95 (June 1998), 6750-6755.
- Jeremy Gunawardena, Chemical reaction network theory for in-silico biologists.
- Martin Feinberg, Lectures on reaction networks.
I should admit that the first two talk about ‘chemical reaction networks’ instead of Petri nets. That’s no big deal: any chemical reaction network gives a Petri net in a pretty obvious way. You can probably figure out how; if you get stuck, just ask.
Also, I should admit that Petri net people say place where I’m saying state.
Here are some other references, which aren’t free unless you have an online subscription or access to a library:
- Peter J. Haas, Stochastic Petri Nets: Modelling, Stability, Simulation, Springer, Berlin, 2002.
- F. Horn and R. Jackson, General mass action kinetics, Archive for Rational Mechanics and Analysis 47 (1972), 81–116.
- Ina Koch, Petri nets – a mathematical formalism to analyze chemical reaction networks, Molecular Informatics 29 (2010), 838–843.
- Darren James Wilkinson, Stochastic Modelling for Systems Biology, Taylor & Francis, New York, 2006.
biology, mathematics, networks, physics. You can follow any responses to this entry through the RSS 2.0 feed. You can leave a response, or trackback from your own site.
I’ve been invited to join this organization. But you can join too:
I hadn’t heard of it before. Do you know anything about it? Here’s their mission statement:
The Lifeboat Foundation is a nonprofit nongovernmental organization dedicated to encouraging scientific advancements while helping humanity survive existential risks and possible misuse of increasingly powerful technologies, including genetic engineering, nanotechnology, and robotics/AI, as we move towards the Singularity.Lifeboat Foundation is pursuing a variety of options, including helping to accelerate the development of technologies to defend humanity, including new methods to combat viruses (such as RNA interference and new vaccine methods), effective nanotechnological defensive strategies, and even self-sustaining spacecolonies in case the other defensive strategies fail.We believe that, in some situations, it might be feasible to relinquish technological capacity in the public interest (for example, we are against the U.S. government posting the recipe for the 1918 flu virus on the Internet).
It seems to have Nick Bostrom and Ray Kurzweil as two of its guiding figures: the overview features quotes from both.
OverviewAn existential risk is a risk that is both global and terminal. Nick Bostrom defines it as a risk “where an adverse outcome would either annihilate Earth-originating intelligent life or permanently and drastically curtail its potential”. The term is frequently used to describe disaster and doomsday scenarios caused by non-friendly superintelligence, misuse of molecular nanotechnology, or other sources of danger.The Lifeboat Foundation was formed to prevent existential events from happening, as once they occur, humanity may have no possibility to correct the error. Unfortunately governments, and humanity in general, always react AFTER a disaster has happened, and some disasters will leave no survivors so we must react BEFORE they occur. We must be proactive.The Lifeboat Foundation is developing programs to prevent existential events (“shields”) as well as programs to preserve civilization (“preservers”) to survive such events.Quotes“Our approach to existential risks cannot be one of trial-and-error. There is no opportunity to learn from errors. The reactive approach — see what happens, limit damages, and learn from experience — is unworkable. Rather, we must take a proactive approach. This requires foresight to anticipate new types of threats and a willingness to take decisive preventive action and to bear the costs (moral and economic) of such actions.” — Nick Bostrom“We cannot rely on trial-and-error approaches to deal with existential risks… We need to vastly increase our investment in developing specific defensive technologies… We are at the critical stage today for biotechnology, and we will reach the stage where we need to directly implement defensive technologies for nanotechnology during the late teen years of this century… A self-replicating pathogen, whether biological or nanotechnology based, could destroy our civilization in a matter of days or weeks.” — Ray Kurzweil
You’ll note there’s no mention here of global warming, mass extinction of species, oil depletion and other minor nuisances. Some people consider these problems insufficiently severe to count as “existential threats”… and thus, perhaps, best left to others. Some argue that there are already enough people worrying about these problem—while other threats need more attention than they’re getting.
That would be an interesting discussion to have. But I’m afraid there’s a cultural divide between the “green crowd” and the “tech crowd” that hinders such a discussion. The green crowd worries about things like global warming, the mass extinction that may currently be underway, and peak oil. The tech crowd worries about things like nanotechnology, artificial intelligence and asteroids hitting the Earth. Each crowd tends to think the other is a bit silly… and they don’t talk to each other enough. Am I just imagining this? I don’t think so.
Of course, any generalization this vast admits many exceptions. I like Gregory Benford because he confounds naive expectations: he thinks global warming is a desperately urgent problem that overshadows all others, but he’s willing to contemplate high-tech solutions. According to my theory, that should annoy both the green crowd and the tech crowd.
Personally I think all significant threats to civilization and biosphere should be evaluated and addressed in a unified way. Setting some aside because they’re “non-existential” or overly studied seems just as dangerous as setting others aside because they seem improbable or science-fiction-esque.
For one thing, I can imagine scenarios where medium-sized problems snowball into big “existential” ones. What’s the chance that in this century, global warming leads to droughts and famines which combined with oil shortages lead to political instability, the collapse of democratic governments, wars… and finally a world-wide nuclear or biological war? Maybe low… but I bet it’s higher than the chance of an asteroid hitting the Earth in this century.
I’m pleased to see that the Lifeboat Foundation plans “future programs” that will appeal to the green crowd:
However, their current programs are strongly focused on issues that appeal to the tech crowd. Maybe that’s okay, but maybe it’s a bit unbalanced:
To protect against unfriendly AI (Artificial Intelligence).To protect against devastating asteroid strikes.To protect against bioweapons and pandemics.As the Internet grows in importance, an attack on it could cause physical as well as informational damage. An attack today on hospital systems or electric utilities could lead to deaths. In the future an attack could be used to alter the output that is produced bynanofactories worldwide leading to massive deaths.Developing fallback positions on Earth in case programs such as our BioShield and NanoShield fail globally or locally.To protect against ecophages and nonreplicatingnanoweapons.This shield strives to protect scientists from obstacles that would prevent latter day Max Plancks from completing their research.To prevent nuclear, biological, and nanotechnological attacks from occurring by using surveillance and sousveillance to identify terrorists before they are able to launch their attacks.To build fail-safes against global existential risks by encouraging the spread of sustainable human civilization beyond Earth.
As we saw last time, a Petri net is a picture that shows different kinds of things and processes that turn bunches of things into other bunches of things, like this:
The kinds of things are called states and the processes are called transitions. We see such transitions in chemistry:
H + OH → H2O
and population biology:
amoeba → amoeba + amoeba
and the study of infectious diseases:
infected + susceptible → infected + infected
and many other situations.
A “stochastic” Petri net says the rate at which each transition occurs. We can think of these transitions as occurring randomly at a certain rate—and then we get a stochastic process described by something called the “master equation”. But for starters, we’ve been thinking about the limit where there are very many things in each state. Then the randomness washes out, and the expected number of things in each state changes deterministically in a manner described by the “rate equation”.
It’s time to explain the general recipe for getting this rate equation! It looks complicated at first glance, so I’ll briefly state it, then illustrate it with tons of examples, and then state it again.
One nice thing about stochastic Petri nets is that they let you dabble in many sciences. Last time we got a tiny taste of how they show up in population biology. This time we’ll look at chemistry and models of infectious diseases. I won’t dig very deep, but take my word for it: you can do a lot with stochastic Petri nets in these subjects! I’ll give some references in case you want to learn more.
Rate equations: the general recipe
Here’s the recipe, really quick:
A stochastic Petri net has a set of states and a set of transitions. Let’s concentrate our attention on a particular transition. Then the th state will appear times as the input to that transition, and times as the output. Our transition also has a reaction rate .
The rate equation answers this question:
where is the number of things in the th state at time . The answer is a sum of terms, one for each transition. Each term works the same way. For the transition we’re looking at, it’s
The factor of shows up because our transition destroys things in the th state and creates of them. The big product over all states, , shows up because our transition occurs at a rate proportional to the product of the numbers of things it takes as inputs. The constant of proportionality is the reaction rate .
The formation of water (1)
But let’s do an example. Here’s a naive model for the formation of water from atomic hydrogen and oxygen:

This Petri net has just one transition: two hydrogen atoms and an oxygen atom collide simultaneously and form a molecule of water. That’s not really how it goes… but what if it were? Let’s use for the number of hydrogen atoms, and so on, and let the reaction rate be . Then we get this rate equation:
\begin{array}{ccl} \frac{d [\mathrm{H}]}{d t} &=& - 2 \alpha [\mathrm{H}]^2 [\mathrm{O}] \\ \\ \frac{d [\mathrm{O}]}{d t} &=& - \alpha [\mathrm{H}]^2 [\mathrm{O}] \\ \\ \frac{d [\mathrm{H}_2\mathrm{O}]}{d t} &=& \alpha [\mathrm{H}]^2 [\mathrm{O}] \end{array}
See how it works? The reaction occurs at a rate proportional to the product of the numbers of things that appear as inputs: two H’s and one O. The constant of proportionality is the rate constant . So, the reaction occurs at a rate equal to . Then:
- Since two hydrogen atoms get used up in this reaction, we get a factor of in the first equation.
- Since one oxygen atom gets used up, we get a factor of in the second equation.
- Since one water molecule is formed, we get a factor of in the third equation.
The formation of water (2)
Let me do another example, just so chemists don’t think I’m an absolute ninny. Chemical reactions rarely proceed by having three things collide simultaneously—it’s too unlikely. So, for the formation of water from atomic hydrogen and oxygen, there will typically be an intermediate step. Maybe something like this:

Here OH is called a ‘hydroxyl radical’. I’m not sure this is the most likely pathway, but never mind—it’s a good excuse to work out another rate equation. If the first reaction has rate constant and the second has rate constant , here’s what we get:
\begin{array}{ccl} \frac{d [\mathrm{H}]}{d t} &=& - \alpha [\mathrm{H}] [\mathrm{O}] - \beta [\mathrm{H}] [\mathrm{OH}] \\ \\ \frac{d [\mathrm{OH}]}{d t} &=& \alpha [\mathrm{H}] [\mathrm{O}] - \beta [\mathrm{H}] [\mathrm{OH}] \\ \\ \frac{d [\mathrm{O}]}{d t} &=& - \alpha [\mathrm{H}] [\mathrm{O}] \\ \\ \frac{d [\mathrm{H}_2\mathrm{O}]}{d t} &=& \beta [\mathrm{H}] [\mathrm{OH}] \end{array}
See how it works? Each reaction occurs at a rate proportional to the product of the numbers of things that appear as inputs. We get minus signs when a reaction destroys one thing of a given kind, and plus signs when it creates one. We don’t get factors of 2 as we did last time, because now no reaction creates or destroys two of anything.
The dissociation of water (1)
In chemistry every reaction comes with a reverse reaction. So, if hydrogen and oxygen atoms can combine to form water, a water molecule can also ‘dissociate’ into hydrogen and oxygen atoms. The rate constants for the reverse reaction can be different than for the original reaction… and all these rate constants depend on the temperature. At room temperature, the rate constant for hydrogen and oxygen to form water is a lot higher than the rate constant for the reverse reaction. That’s why we see a lot of water, and not many lone hydrogen or oxygen atoms. But at sufficiently high temperatures, the rate constants change, and water molecules become more eager to dissociate.
Calculating these rate constants is a big subject. I’m just starting to read this book, which looked like the easiest one on the library shelf:
- S. R. Logan, Chemical Reaction Kinetics, Longman, Essex, 1996.
But let’s not delve into these mysteries today. Let’s just take our naive Petri net for the formation of water and turn around all the arrows, to get the reverse reaction:

If the reaction rate is , here’s the rate equation:
\begin{array}{ccl} \frac{d [\mathrm{H}]}{d t} &=& 2 \alpha [\mathrm{H}_2\mathrm{O}] \\ \\ \frac{d [\mathrm{O}]}{d t} &=& \alpha [\mathrm{H}_2 \mathrm{O}] \\ \\ \frac{d [\mathrm{H}_2\mathrm{O}]}{d t} &=& - \alpha [\mathrm{H}_2 \mathrm{O}] \end{array}
See how it works? The reaction occurs at a rate proportional to , since it has just a single water molecule as input. That’s where the comes from. Then:
- Since two hydrogen atoms get formed in this reaction, we get a factor of +2 in the first equation.
- Since one oxygen atom gets formed, we get a factor of +1 in the second equation.
- Since one water molecule gets used up, we get a factor of +1 in the third equation.
The dissociation of water (part 2)
Of course, we can also look at the reverse of the more realistic reaction involving a hydroxyl radical as an intermediate. Again, we just turn around the arrows in the Petri net we had:

Now the rate equation looks like this:
\begin{array}{ccl} \frac{d [\mathrm{H}]}{d t} &=& + \alpha [\mathrm{OH}] + \beta [\mathrm{H}_2\mathrm{O}] \\ \\ \frac{d [\mathrm{OH}]}{d t} &=& - \alpha [\mathrm{OH}] + \beta [\mathrm{H}_2 \mathrm{O}] \\ \\ \frac{d [\mathrm{O}]}{d t} &=& + \alpha [\mathrm{OH}] \\ \\ \frac{d [\mathrm{H}_2\mathrm{O}]}{d t} &=& - \beta [\mathrm{H}_2\mathrm{O}] \end{array}
Do you see why? Test your understanding of the general recipe.
By the way: if you’re a category theorist, when I said “turn around all the arrows” you probably thought “opposite category”. And you’d be right! A Petri net is just a way of presenting a strict symmetric monoidal category that’s freely generated by some objects (the states) and some morphisms (the transitions). When we turn around all the arrows in our Petri net, we’re getting a presentation of the opposite symmetric monoidal category. For more details, try:
- Vladimiro Sassone, On the category of Petri net computations, 6th International Conference on Theory and Practice of Software Development, Proceedings of TAPSOFT ’95, Lecture Notes in Computer Science 915, Springer, Berlin, pp. 334-348.
After I explain how stochastic Petri nets are related to quantum field theory, I hope to say more about this category theory business. But if you don’t understand it, don’t worry about it now—let’s do a few more examples.
The SI model
The SI model is an extremely simple model of an infectious disease. We can describe it using this Petri net:

There are two states: susceptible and infected. And there’s a transition called infection, where an infected person meets a susceptible person and infects them.
Suppose is the number of susceptible people and the number of infected ones. If the rate constant for infection is the rate equation is
Do you see why?
By the way, it’s easy to solve these equations exactly. The total number of people doesn’t change, so is a conserved quantity. Use this to get rid of one of the variables. You’ll get a version of the famous logistic equation, so the fraction of people infected must grow sort of like this:

Puzzle. Is there a stochastic Petri net with just one state whose rate equation is the logistic equation:
The SIR model
The SI model is just a warmup for the more interesting SIR model, which was invented by Kermack and McKendrick in 1927:
- W. O. Kermack and A. G. McKendrick, A Contribution to the mathematical theory of epidemics, Proc. Roy. Soc. Lond. A 115 (1927), 700-721.
The SIR model has an extra state, called resistant, and an extra transition, called recovery, where an infected person gets better and develops resistance to the disease:

If the rate constant for infection is and the rate constant for recovery is , the rate equation for this stochastic Petri net is:
See why?
I don’t know a ‘closed-form’ solution to these equations. But Kermack and McKendrick found an approximate solution in their original paper. They used this to model the death rate from bubonic plague during an outbreak in Bombay that started in 1905, and they got pretty good agreement. Nowadays, of course, we can solve these equations numerically on the computer.
The SIRS model
There’s an even more interesting model of infectious disease called the SIRS model. This has one more transition, called losing resistance, where a resistant person can go back to being susceptible. Here’s the Petri net:

Puzzle. If the rate constants for recovery, infection and loss of resistance are and , write down the rate equations for this stochastic Petri net.
In the SIRS model we see something new: cyclic behavior! Say you start with a few infected people and a lot of susceptible ones. Then lots of people get infected… then lots get resistant… and then, much later, if you set the rate constants right, they lose their resistance and they’re ready to get sick all over again! You can sort of see it from the Petri net, which looks like a cycle.
I learned about the SI, SIR and SIRS models from this great book:
- Marc Mangel, The Theoretical Biologist’s Toolbox: Quantitative Methods for Ecology and Evolutionary Biology, Cambridge U. Press, Cambridge, 2006.
For more models of this type, see:
- Compartmental models in epidemiology, Wikipedia.
A ‘compartmental model’ is closely related to a stochastic Petri net, but beware: the pictures in this article are not really Petri nets!
The general recipe revisited
Now let me remind you of the general recipe and polish it up a bit. So, suppose we have a stochastic Petri net with states. Let be the number of things in the th state. Then the rate equation looks like:
It’s really a bunch of equations, one for each . But what is the right-hand side?
The right-hand side is a sum of terms, one for each transition in our Petri net. So, let’s assume our Petri net has just one transition! (If there are more, consider one at a time, and add up the results.)
Suppose the th state appears as input to this transition times, and as output times. Then the rate equation is
where is the rate constant for this transition.
That’s really all there is to it! But subscripts make my eyes hurt more and more as I get older—this is the real reason for using index-free notation, despite any sophisticated rationales you may have heard—so let’s define a vector
that keeps track of how many things there are in each state. Similarly let’s make up an input vector:
and an output vector:
for our transition. And a bit more unconventionally, let’s define
Then we can write the rate equation for a single transition as
This looks a lot nicer!
Indeed, this emboldens me to consider a general stochastic Petri net with lots of transitions, each with their own rate constant. Let’s write for the set of transitions and for the rate constant of the transition . Let and be the input and output vectors of the transition . Then the rate equation for our stochastic Petri net is
That’s the fully general recipe in a nutshell. I’m not sure yet how helpful this notation will be, but it’s here whenever we want it.
Next time we’ll get to the really interesting part, where ideas from quantum theory enter the game! We’ll see how things in different states randomly transform into each other via the transitions in our Petri net. And someday we’ll check that the expected number of things in each state evolves according to the rate equation we just wrote down… at least in there limit where there are lots of things in each state.
chemistry, networks. You can follow any responses to this entry through the RSS 2.0 feed. You can leave a response, or trackback from your own site.
Last time I explained the rate equation of a stochastic Petri net. But now let’s get serious: let’s see what’s really stochastic— that is, random—about a stochastic Petri net. For this we need to forget the rate equation (temporarily) and learn about the ‘master equation’. This is where ideas from quantum field theory start showing up!
A Petri net has a bunch of states and a bunch of transitions. Here’s an example we’ve already seen, from chemistry:
The states are in yellow, the transitions in blue. A labelling of our Petri net is a way of putting some number of things in each state. We can draw these things as little black dots:
In this example there are only 0 or 1 things in each state: we’ve got one atom of carbon, one molecule of oxygen, one molecule of sodium hydroxide, one molecule of hydrochloric acid, and nothing else. But in general, we can have any natural number of things in each state.
In a stochastic Petri net, the transitions occur randomly as time passes. For example, as time passes we could see a sequence of transitions like this:




Each time a transition occurs, the number of things in each state changes in an obvious way.
The Master Equation
Now, I said the transitions occur ‘randomly’, but that doesn’t mean there’s no rhyme or reason to them! The miracle of probability theory is that it lets us state precise laws about random events. The law governing the random behavior of a stochastic Petri net is called the ‘master equation’.
In a stochastic Petri net, each transition has a rate constant, a nonnegative real number. Roughly speaking, this determines the probability of that transition.
A bit more precisely: suppose we have a Petri net that is labelled in some way at some moment. Then the probability that a given transition occurs in a short time is approximately:
- the rate constant for that transition, times
- the time , times
- the number of ways the transition can occur.
More precisely still: this formula is correct up to terms of order . So, taking the limit as , we get a differential equation describing precisely how the probability of the Petri net having a given labelling changes with time! And this is the master equation.
Now, you might be impatient to actually see the master equation, but that would be rash. The true master doesn’t need to see the master equation. It sounds like a Zen proverb, but it’s true. The raw beginner in mathematics wants to see the solutions of an equation. The more advanced student is content to prove that the solution exists. But the master is content to prove that the equation exists.
A bit more seriously: what matters is understanding the rules that inevitably lead to some equation: actually writing it down is then straightforward.
And you see, there’s something I haven’t explained yet: “the number of ways the transition can occur”. This involves a bit of counting. Consider, for example, this Petri net:
Suppose there are 10 rabbits and 5 wolves.
- How many ways can the birth transition occur? Since birth takes one rabbit as input, it can occur in 10 ways.
- How many ways can predation occur? Since predation takes one rabbit and one wolf as inputs, it can occur in 10 × 5 = 50 ways.
- How many ways can death occur? Since death takes one wolf as input, it can occur in 5 ways.
Or consider this one:
Suppose there are 10 hydrogen atoms and 5 oxygen atoms. How many ways can they form a water molecule? There are 10 ways to pick the first hydrogen, 9 ways to pick the second hydrogen, and 5 ways to pick the oxygen. So, there are
ways.
Note that we’re treating the hydrogen atoms as distinguishable, so there are ways to pick them, not . In general, the number of ways to choose distinguishable things from a collection of is the falling power
where there are factors in the product, but each is 1 less than the preceding one—hence the term ‘falling’.
Okay, now I’ve given you all the raw ingredients to work out the master equation for any stochastic Petri net. The previous paragraph was a big fat hint. One more nudge and you’re on your own:
Puzzle. Suppose we have a stochastic Petri net with states and just one transition, whose rate constant is . Suppose the th state appears times as the input of this transition and times as the output. A labelling of this stochastic Petri net is a -tuple of natural numbers saying how many things are in each state. Let be the probability that the labelling is at time . Then the master equation looks like this:
for some matrix of real numbers What is this matrix?
You can write down a formula for this matrix using what I’ve told you. And then, if you have a stochastic Petri net with more transitions, you can just compute the matrix for each transition using this formula, and add them all up.
Someday I’ll tell you the answer to this puzzle, but I want to get there by a strange route: I want to guess the master equation using ideas from quantum field theory!
Some clues
Why?
Well, if we think about a stochastic Petri net whose labelling undergoes random transitions as I’ve described, you’ll see that any possible ‘history’ for the labelling can be drawn in a way that looks like a Feynman diagram. In quantum field theory, Feynman diagrams show how things interact and turn into other things. But that’s what stochastic Petri nets do, too!
For example, if our Petri net looks like this:
then a typical history can be drawn like this:

Some rabbits and wolves come in on top. They undergo some transitions as time passes, and go out on the bottom. The vertical coordinate is time, while the horizontal coordinate doesn’t really mean anything: it just makes the diagram easier to draw.
If we ignore all the artistry that makes it cute, this Feynman diagram is just a graph with states as edges and transitions as vertices. Each transition occurs at a specific time.
We can use these Feynman diagrams to compute the probability that if we start it off with some labelling at time , our stochastic Petri net will wind up with some other labelling at time . To do this, we just take a sum over Feynman diagrams that start and end with the given labellings. For each Feynman diagram, we integrate over all possible times at which the transitions occur. And what do we integrate? Just the product of the rate constants for those transitions!
That was a bit of a mouthful, and it doesn’t really matter if you followed it in detail. What matters is that it sounds a lot like stuff you learn when you study quantum field theory!
That’s one clue that something cool is going on here. Another is the master equation itself:
This looks a lot like Schrödinger’s equation, the basic equation describing how a quantum system changes with the passage of time.
We can make it look even more like Schrödinger’s equation if we create a vector space with the labellings as a basis. The numbers will be the components of some vector in this vector space. The numbers will be the matrix entries of some operator on that vector space. And the master equation becomes:
Compare Schrödinger’s equation:
The only visible difference is that factor of !
But of course this is linked to another big difference: in the master equation describes probabilities, so it’s a vector in a real vector space. In quantum theory describes probability amplitudes, so it’s a vector in a complex Hilbert space.
Apart from this huge difference, everything is a lot like quantum field theory. In particular, our vector space is a lot like the Fock space one sees in quantum field theory. Suppose we have a quantum particle that can be in different states. Then its Fock space is the Hilbert space we use to describe an arbitrary collection of such particles. It has an orthonormal basis denoted
where are natural numbers saying how many particles there are in each state. So, any vector in Fock space looks like this:
But if write the whole list simply as , this becomes
This is almost like what we’ve been doing with Petri nets!—except I hadn’t gotten around to giving names to the basis vectors.
In quantum field theory class, I learned lots of interesting operators on Fock space: annihilation and creation operators, number operators, and so on. So, when I bumped into this master equation
it seemed natural to take the operator and write it in terms of these. There was an obvious first guess, which didn’t quite work… but thinking a bit harder eventually led to the right answer. Later, it turned out people had already thought about similar things. So, I want to explain this.
When I first started working on this stuff, I was focused on the difference between collections of indistinguishable things, like bosons or fermions, and collections of distinguishable things, like rabbits or wolves.
But with the benefit of hindsight, it’s even more important to think about the difference between ordinary quantum theory, which is all about probability amplitudes, and the game we’re playing now, which is all about probabilities. So, next time I’ll explain how we need to modify quantum theory so that it’s about probabilities. This will make it easier to guess a nice formula for .
mathematics, networks, physics, probability. You can follow any responses to this entry through the RSS 2.0 feed. You can leave a response, or trackback from your own site.
Yes, one of the mischievous trio is cobalt.
In German a goblin is called a “Kobold”. Miners called certain minerals “Kobold ore”, or goblin ore, because they were poor in known metals and gave poisonous arsenic-containing fumes when smelted. In 1735, such ores were found to contain a new metal – the first discovered since antiquity – and this metal was called cobalt
I don’t think of either Niobe or Athena as a “mischievous sprite”.
From the Wikipedia article Kobold:
The kobold (or kobolt) is a sprite stemming from Germanic mythology and surviving into modern times in German folklore. Although usually invisible, a kobold can materialise in the form of an animal, fire, a human being, and a mundane object. The most common depictions of kobolds show them as humanlike figures the size of small children. Kobolds who live in human homes wear the clothing of peasants; those who live in mines are hunched and ugly; and kobolds who live on ships smoke pipes and wear sailor clothing.Legends tell of three major types of kobolds. Most commonly, the creatures are house spirits of ambivalent nature; while they sometimes perform domestic chores, they play malicious tricks if insulted or neglected. Famous kobolds of this type include King Goldemar, Heinzelmann, Hödekin. In some regions, kobolds are known by local names, such as the Galgenmännlein of southern Germany and the Heinzelmännchen of Cologne. Another type of kobold haunts underground places, such as mines. The name of the element cobalt comes from the creature’s name, because medieval miners blamed the sprite for the poisonous and troublesome nature of the typical arsenical ores of this metal (cobaltite and smaltite) which polluted other mined elements. A third kind of kobold, the Klabautermann, lives aboard ships and helps sailors.
Last time we saw clues that stochastic Petri nets are a lot like quantum field theory, but with probabilities replacing amplitudes. There’s a powerful analogy at work here, which can help us a lot. So, this time I want to make that analogy precise.
But first, let me quickly sketch why it could be worthwhile.
A Poisson process
Consider this stochastic Petri net with rate constant :
It describes an inexhaustible supply of fish swimming down a river, and getting caught when they run into a fisherman’s net. In any short time there’s a chance of about of a fish getting caught. There’s also a chance of two or more fish getting caught, but this becomes negligible by comparison as . Moreover, the chance of a fish getting caught during this interval of time is independent of what happens before or afterwards. This sort of process is called a Poisson process.
Problem. Suppose we start out knowing for sure there are no fish in the fisherman’s net. What’s the probability that he has caught fish at time ?
At any time there will be some probability of having caught fish; let’s call this probability , or for short. We can summarize all these probabilities in a single power series, called a generating function:
Here is a formal variable—don’t ask what it means, for now it’s just a trick. In quantum theory we use this trick when talking about collections of photons rather than fish, but then the numbers are complex ‘amplitudes’. Now they are real probabilities, but we can still copy what the physicists do, and use this trick to rewrite the master equation as follows:
This describes how the probability of having caught any given number of fish changes with time.
What’s the operator ? Well, in quantum theory we describe the creation of photons using a certain operator on power series called the creation operator:
We can try to apply this to our fish. If at some time we’re 100% sure we have fish, we have
so applying the creation operator gives
One more fish! That’s good. So, an obvious wild guess is
where is the rate at which we’re catching fish. Let’s see how well this guess works.
If you know how to exponentiate operators, you know to solve this equation:
It’s easy:
Since we start out knowing there are no fish in the net, we have
so with our guess for we get
But is the operator of multiplication by , so is multiplication by , so
So, if our guess is right, the probability of having caught fish at time is
Unfortunately, this can’t be right, because these probabilities don’t sum to 1! Instead their sum is
We can try to wriggle out of the mess we’re in by dividing our answer by this fudge factor. It sounds like a desperate measure, but we’ve got to try something!
This amounts to guessing that the probability of having caught fish by time is
And this is right! This is called the Poisson distribution: it’s famous for being precisely the answer to the problem we’re facing.
So on the one hand our wild guess about was wrong, but on the other hand it was not so far off. We can fix it as follows:
The extra gives us the fudge factor we need.
So, a wild guess corrected by an ad hoc procedure seems to have worked! But what’s really going on?
What’s really going on is that , or any multiple of this, is not a legitimate Hamiltonian for a master equation: if we define a time evolution operator using a Hamiltonian like this, probabilities won’t sum to 1! But is okay. So, we need to think about which Hamiltonians are okay.
In quantum theory, self-adjoint Hamiltonians are okay. But in probability theory, we need some other kind of Hamiltonian. Let’s figure it out.
Probability versus quantum theory
Suppose we have a system of any kind: physical, chemical, biological, economic, whatever. The system can be in different states. In the simplest sort of model, we say there’s some set of states, and say that at any moment in time the system is definitely in one of these states. But I want to compare two other options:
- In a probabilistic model, we may instead say that the system has a probability of being in any state . These probabilities are nonnegative real numbers with
- In a quantum model, we may instead say that the system has an amplitude of being in any state . These amplitudes are complex numbers with
Probabilities and amplitudes are similar yet strangely different. Of course given an amplitude we can get a probability by taking its absolute value and squaring it. This is a vital bridge from quantum theory to probability theory. Today, however, I don’t want to focus on the bridges, but rather the parallels between these theories.
We often want to replace the sums above by integrals. For that we need to replace our set by a measure space, which is a set equipped with enough structure that you can integrate real or complex functions defined on it. Well, at least you can integrate so-called ‘integrable’ functions—but I’ll neglect all issues of analytical rigor here. Then:
- In a probabilistic model, the system has a probability distribution , which obeys and
- In a quantum model, the system has a wavefunction , which obeys
In probability theory, we integrate over a set to find out the probability that our systems state is in this set. In quantum theory we integrate over the set to answer the same question.
We don’t need to think about sums over sets and integrals over measure spaces separately: there’s a way to make any set into a measure space such that by definition,
In short, integrals are more general than sums! So, I’ll mainly talk about integrals, until the very end.
In probability theory, we want our probability distributions to be vectors in some vector space. Ditto for wave functions in quantum theory! So, we make up some vector spaces:
- In probability theory, the probability distribution is a vector in the space
- In quantum theory, the wavefunction is a vector in the space
You may wonder why I defined to consist of complex functions when probability distributions are real. I’m just struggling to make the analogy seem as strong as possible. In fact probability distributions are not just real but nonnegative. We need to say this somewhere… but we can, if we like, start by saying they’re complex-valued functions, but then whisper that they must in fact be nonnegative (and thus real). It’s not the most elegant solution, but that’s what I’ll do for now.
Now:
- The main thing we can do with elements of , besides what we can do with vectors in any vector space, is integrate one. This gives a linear map:
- The main thing we can with elements of , besides the besides the things we can do with vectors in any vector space, is take the inner product of two:
This gives a map that’s linear in one slot and conjugate-linear in the other:
First came probability theory with ; then came quantum theory with . Naive extrapolation would say it’s about time for someone to invent an even more bizarre theory of reality based on In this, you’d have to integrate the product of three wavefunctions to get a number! The math of Lp spaces is already well-developed, so give it a try if you want. I’ll stick to and today.
Stochastic versus unitary operators
Now let’s think about time evolution:
- In probability theory, the passage of time is described by a map sending probability distributions to probability distributions. This is described using a stochastic operator
meaning a linear operator such that
and
- In quantum theory the passage of time is described by a map sending wavefunction to wavefunctions. This is described using an isometry
meaning a linear operator such that
In quantum theory we usually want time evolution to be reversible, so we focus on isometries that have inverses: these are called unitary operators. In probability theory we often consider stochastic operators that are not invertible.
Infinitesimal stochastic versus self-adjoint operators
Sometimes it’s nice to think of time coming in discrete steps. But in theories where we treat time as continuous, to describe time evolution we usually need to solve a differential equation. This is true in both probability theory and quantum theory:
- In probability theory we often describe time evolution using a differential equation called the master equation:
whose solution is
- In quantum theory we often describe time evolution using a differential equation called Schrödinger’s equation:
whose solution is
In fact the appearance of in the quantum case is purely conventional; we could drop it to make the analogy better, but then we’d have to work with ‘skew-adjoint’ operators instead of self-adjoint ones in what follows.
Let’s guess what properties an operator should have to make unitary for all . We start by assuming it’s an isometry:
Then we differentiate this with respect to and set , getting
or in other words
Physicists call an operator obeying this condition self-adjoint. Mathematicians know there’s more to it, but today is not the day to discuss such subtleties, intriguing though they be. All that matters now is that there is, indeed, a correspondence between self-adjoint operators and well-behaved ‘one-parameter unitary groups’ . This is called Stone’s Theorem.
But now let’s copy this argument to guess what properties an operator must have to make stochastic. We start by assuming is stochastic, so
and
We can differentiate the first equation with respect to and set , getting
for all .
But what about the second condition,
It seems easier to deal with this in the special case when integrals over reduce to sums. So let’s suppose that happens… and let’s start by seeing what the first condition says in this case.
In this case, has a basis of ‘Kronecker delta functions’: The Kronecker delta function vanishes everywhere except at one point , where it equals 1. Using this basis, we can write any operator on as a matrix.
As a warmup, let’s see what it means for an operator
to be stochastic in this case. We’ll take the conditions
and
and rewrite them using matrices. For both, it’s enough to consider the case where is a Kronecker delta, say .
In these terms, the first condition says
for each column . The second says
for all . So, in this case, a stochastic operator is just a square matrix where each column sums to 1 and all the entries are nonnegative. (Such matrices are often called left stochastic.)
Next, let’s see what we need for an operator to have the property that is stochastic for all . It’s enough to assume is very small, which lets us use the approximation
and work to first order in . Saying that each column of this matrix sums to 1 then amounts to
which requires
Saying that each entry is nonnegative amounts to
When this will be automatic when is small enough, so the meat of this condition is
So, let’s say is an infinitesimal stochastic matrix if its columns sum to zero and its off-diagonal entries are nonnegative.
I don’t love this terminology: do you know a better one? There should be some standard term. People here say they’ve seen the term ‘stochastic Hamiltonian’. The idea behind my term is that any infintesimal stochastic operator should be the infinitesimal generator of a stochastic process.
In other words, when we get the details straightened out, any 1-parameter family of stochastic operators
obeying
and continuity:
should be of the form
for a unique ‘infinitesimal stochastic operator’ .
When is a finite set, this is true—and an infinitesimal stochastic operator is just a square matrix whose columns sum to zero and whose off-diagonal entries are nonnegative. But do you know a theorem characterizing infinitesimal stochastic operators for general measure spaces ? Someone must have worked it out.
Luckily, for our work on stochastic Petri nets, we only need to understand the case where is a countable set and our integrals are really just sums. This should be almost like the case where is a finite set—but we’ll need to take care that all our sums converge.
The moral
Now we can see why a Hamiltonian like is no good, while is good. (I’ll ignore the rate constant since it’s irrelevant here.) The first one is not infinitesimal stochastic, while the second one is!
In this example, our set of states is the natural numbers:
The probability distribution
tells us the probability of having caught any specific number of fish.
The creation operator is not infinitesimal stochastic: in fact, it’s stochastic! Why? Well, when we apply the creation operator, what was the probability of having fish now becomes the probability of having fish. So, the probabilities remain nonnegative, and their sum over all is unchanged. Those two conditions are all we need for a stochastic operator.
Using our fancy abstract notation, these conditions say:
and
So, precisely by virtue of being stochastic, the creation operator fails to be infinitesimal stochastic:
Thus it’s a bad Hamiltonian for our stochastic Petri net.
On the other hand, is infinitesimal stochastic. Its off-diagonal entries are the same as those of , so they’re nonnegative. Moreover:
precisely because
You may be thinking: all this fancy math just to understand a single stochastic Petri net, the simplest one of all!

But next time I’ll explain a general recipe which will let you write down the Hamiltonian for any stochastic Petri net. The lessons we’ve learned today will make this much easier. And pondering the analogy between probability theory and quantum theory will also be good for our bigger project of unifying the applications of network diagrams to dozens of different subjects.
mathematics, networks, physics. You can follow any responses to this entry through the RSS 2.0 feed. You can leave a response, or trackback from your own site.