Why choreographic programming?¶
Short version: distributed programs are hard because the hard part is the interaction, and choreographic programming lets you write the interaction once and be done with it.
The problem with hand-written communication¶
When you write one program per role in the classic way, every protocol lives implicitly in the interleaving of three programs. The bugs live in the places no single file shows you:
- Deadlock — role A waits for a message that role B never sends because a guard evaluated differently at each role.
- Race on state — two roles both "update the counter" and the agreed value silently diverges.
- Drift — you change the protocol in A's code but forget B's, and the system fails only under a rare interleaving.
These are exactly the bugs that are hardest to reproduce, hardest to test, and most expensive in production. They are global bugs living in local code.
What a choreography gives you¶
A choreography is the whole interaction written as one program:
- One source of truth. The protocol appears exactly once. Change the election, the aggregate, the choreography — every role follows automatically.
- Correct communication by construction. The generated per-role programs
contain exactly the sends and receives the choreography names. There is no
"message nobody listens to": every
move/copynames a concrete source and destination, matched by construction. - No deadlock, no divergence. Because every send has a matching receive, and every branch/loop is decided according to knowledge of choice (only roles that can see the choice participate, or are explicitly told), the projections terminate together or not at all.
- Location-aware types. A value can't be used where it isn't located. Move
xtoBand try to use it atA? Rejected at definition time. This is the distinctive static check of choreographic programming, and it's the single highest-leverage safety property here: most distributed bugs start with a value being read where it was never delivered.
What it feels like to develop¶
- Write the interaction once (
@choreography def ...), with roles, value locations, and choice spelled out. - Simulate every day.
simulate_chorruns all roles in one process over an in-memory transport — no deployment, no network, instant feedback, and the same code runs the same way. - Swap the transport when you deploy. The projected roles run unchanged
over the bundled
TcpTransport, or any transport you provide (play_role's{role, send, recv, locators}config). The protocol you debugged on your laptop is the protocol you ship.
Why Python makes this practical¶
Klor (Clojure) shows the idea works with elegant macros; Python has no macros, but it doesn't need them:
inspect+astgive the decorator the source it needs to analyze and project — the "compiler" is a normal Python pass.asynciogives the simulator (one task per role), the TCP transport, and the whole network ecosystem for free.- The DSL is a subset of real Python: role annotations,
if/match/while/forare the real statements, so tooling, debugging, and teaching stay familiar.
Where it fits — and where it doesn't¶
Fits: deterministic cooperative protocols and algorithms — leader election, aggregation/combine, broadcast, orchestrated multi-party workflows, remote procedure call, distributed state machines where the steps are known.
Doesn't fit (honestly):
- Failure and adversarial behavior. A choreography assumes every role executes the agreed interaction. There is no crash model, no retry, no Byzantine handling — a crashed or misbehaving role stops the choreography. This is the classic trade-off of choreographic programming (it trades resilience analysis for correctness of the protocol itself).
- Choreography-internal non-determinism. Choice is explicit and visible
(
ifguards,match), and everything else agrees by construction. If you need genuine independence within a step, that belongs inside a role's lifted computation — not in the choreography.
For verifying what the protocol is supposed to be, choreographic programming is a complement to model checking (TLA+, mCRL2): the choreography is a readable, executable, type-checked statement of the protocol; Klor itself was developed alongside the mCRL2 model checker for exactly this reason.
The research lineage¶
KlorPy is a working port of Klor (Clojure), which in turn draws on the theory of choreographic programming and its relatives (multi-party session types, endpoint projection) — see Relationship to Klor and the original Klor papers for the deeper story. This site is the how; that work is the why-well-founded.