Arboria Benchmarks
Two suites, one discipline. The swarm benchmark ranks coordination policies across a breadth of objectives. The Arboria Constellation Benchmark scores them against real orbital geometry — propagated orbits, inter-satellite visibility, eclipse, and contact plans — rather than against an assumed always-on network.
Both exist for the same reason: to separate “we tuned flocking” from “we made measurable progress against prior art.”
Availability. Both suites currently ship inside Gossamer and Orrery, which are proprietary. A versioned, publicly runnable release of the constellation benchmark — with standardised fault models and a leaderboard harness that requires no access to our internal stack — is the deliverable we consider most important, for the reason given at the bottom of this page. Until it lands, treat this as a description of the evaluation protocol rather than something you can execute today. See Reproducibility and Data Availability.
The swarm benchmark
Scenarios
Each scenario fixes an initial state, a per-step reward, and a terminal metric.
| Name | Question | Terminal metric |
|---|---|---|
dispersal | How fast can a clumped swarm spread without colliding? | Mean nearest-neighbour distance at termination |
rendezvous | How fast does a scattered swarm meet at a common point? | Final mean distance to centroid (lower is better) |
coverage | Explore a bounded region; maximize cells visited per unit time | Unique cells visited over total cells |
leader_follower | One agent is exogenously driven; keep the swarm within range | Mean follower distance to the leader’s path |
byzantine | Inject a fraction of actively adversarial agents | Terminal metric of the base scenario, under perturbation |
vicsek_transition | Does the substrate reproduce the known ordering transition? | Polarisation at termination |
Benchmark scenarios are a different thing from Gossamer’s tasks, and the distinction matters. Scenarios exist to rank policies across a breadth of objectives. Tasks exist to define the coordination quality that the delay papers measure, and each is scored against a peer-derived target so that coordination is provably necessary to score well. Scenarios are for comparison; tasks are for measurement.
vicsek_transition is there as an external anchor: it reproduces a transition
that is established in the literature, so a substrate that cannot order at all
fails visibly rather than silently producing plausible numbers for everything else.
Baselines
Every new policy reports against the same reference set, so a reviewer has a stable point of comparison rather than a number in isolation.
random applies uniform random accelerations. It is the lower bound, and a
surprising number of published policies fail to clear it convincingly on at least
one scenario.
greedy is a per-scenario hand-crafted heuristic — move toward the centroid
for rendezvous, push away from the nearest neighbour for dispersal, a persistent
random walk for coverage. It is what a competent engineer writes in an afternoon,
and beating it is the real bar.
gossamer_flocking is classical
Boids, included because it is the
field’s shared reference point.
vicsek_alignment is constant-speed, heading-only alignment — the right
instrument when the question is order, which cohesive flocking is not.
Learned policies — MAPPO and its relatives — are trained and evaluated against the same scenarios through Gossamer’s graph substrate, which deliberately presents classical and learned policies behind one interface so a comparison between them is not confounded by plumbing.
Fault models, and why they have to prove they fired
Faults are the easiest place in a benchmark to ship something that does nothing, because every way a fault row can fail looks like a result. A fault that never triggers, a fault the engine silently ignores, and a fault row run on a substrate with no fault module all produce the same clean leaderboard with every submission tied — and all three read as robustness.
We know this because we shipped it. An early byzantine scenario computed which
agents were adversarial and then never read the marks, so the “byzantine”
leaderboard row was a plain rendezvous run under a different label. Nothing
crashed and nothing looked wrong.
So the evidence is now structural rather than a matter of care. Every fault model names the engine counter that is positive if and only if the fault actually occurred, the harness checks that counter before a result is constructed, and a zero raises instead of returning a row. A fault model with nothing to point at cannot be added to the registry, and a fault row cannot be run on a substrate that has no fault module.
The transient-upset models form a ladder over bit ranges, and the spacing is itself the finding. Flipping a low mantissa bit of a double perturbs it by essentially nothing; flipping an exponent bit is catastrophic. Measured at an identical upset rate and an identical number of flips, the damage spans roughly two orders of magnitude across that ladder. Read the uniform-mantissa row alone and you would conclude a swarm is radiation-robust; the conclusion rests entirely on the exponent bits being protected. That is a quantitative argument about where error correction is worth spending, and it is why a row scoring near the control is the measurement rather than a broken test.
We also keep two failure classes deliberately separate. A hard fault is permanent and marked: the agent stops and its peers can route around a corpse. A single-event upset is transient and silent: one bit of one position or velocity word flips, nothing is marked, and the agent carries on confidently wrong. The second is the one that matters in orbit, and a permanent-failure model cannot express it.
Harness and substrate
The harness runs a policy against a scenario for a fixed step budget across several seeds and emits one row per run: scenario, baseline, agent count, steps, seed, terminal metric, mean reward, elapsed wall-clock, and the substrate it ran on.
That last field matters. The default substrate is the compiled Leviathan engine — the same one the papers run on — because a benchmark that cannot run on the engine the papers use cannot be the neutral standard it exists to be. A pure-NumPy reference stepper exists and can be requested explicitly, but the harness refuses to fall back to it on its own: a benchmark that silently ran on a different substrate than you asked for is not a benchmark, and the number it produces carries no trace of the swap. Results from the two substrates are never mixed within a table.
Reproducibility
Seeds are recorded with each row, and rerunning at the same seed yields identical numbers — Gossamer takes an explicit random generator at every stochastic entry point and makes no module-level random calls anywhere.
The Arboria Constellation Benchmark
The swarm suite asks whether a policy coordinates. The constellation benchmark asks whether it coordinates on a sky that actually exists.
Every task is built on Orrery: orbits are propagated, contact windows are computed from geometry, and the delay a submission experiences is the one that constellation imposes — not a number chosen to make a point. Tasks consume the coordination algorithms lazily and the geometry directly, so every submission is scored against the same sky. That is the entire point of it.
Versioning is the load-bearing part
A benchmark is only a standard if two people quoting it mean the same thing, so:
- The suite carries a version, and results are comparable only within one version.
- Every task exposes a specification digest — a hash of exactly what it ran. A reported number names its configuration rather than trusting a description of it.
- Changing a frozen parameter is a new version, never a knob. There is no flag that quietly alters what a published number meant.
- A modelling choice that could go either way is a specification field, not a runtime option, so two choices are two experiments with two digests rather than one number with an asterisk.
That last rule has teeth. A satellite’s delay is a distribution, and collapsing it to a scalar is a modelling decision — mean, median and 99th percentile are genuinely different experiments, and which one an outcome follows is an open research question we are actively pursuing. So the reduction statistic is part of the spec and part of the digest.
The tasks
orbital-market — decentralised scheduling of compute jobs across a
constellation under eclipse, thermal and downlink constraints. Submissions are
scheduler callables.
contact-plan-aoi — how stale is each satellite’s picture of the others,
given a real contact plan? Submissions are relay policies. There is no
randomness anywhere in this task: a task-and-policy pair is the experiment.
It is anchored to the delay oracle in both directions — flooding must land just
above the earliest-arrival floor, because a submission that beats the physics has
a bug rather than a breakthrough.
dcc-delay — our delay-coordination axis with the knob removed.
Per-satellite staleness comes from propagated geometry over a real constellation,
so a submission is scored against the delay the sky imposes. It runs against both
a designed Walker constellation and a set of real Starlink elements, and the
two disagree in a way that is itself instructive: their mean delays agree to
within 7% while their 90th percentiles differ by more than three orders of
magnitude. A study reporting a scalar mean would call those two constellations
equivalent. They are nothing alike.
Why we want someone else to run it
A benchmark owned by one lab and runnable only by that lab is marketing. The value of this one — to us and to anyone else working on orbital coordination — depends entirely on other people being able to execute it, disagree with our numbers, and submit their own. That is why the public release is the deliverable we care most about, and why the versioning discipline above is not bureaucracy: it is what makes a leaderboard from four different groups mean something.