A Decision Model as a Reflex: A 10-Drone Simulation With and Without a Safety Backstop
Can a decision model acting as a reflex manage traffic and battery for 10 drones on a shared map? A 30-minute experiment with and without a deterministic safety backstop: 12 collisions without it, none with it.

Summary: the experiment and what it tests
The question. Can a fast decision model, Jev, acting as a reflex, manage traffic and low battery for a fleet of drones that share one map, the same waypoints and one cruise altitude (45 m), while keeping crashes low and missions high? Jev does not fly the drones. Roughly once a second, per drone, it answers a choice (CONTINUE, HOLD, SLOW, CHANGE_LEVEL, RETURN_HOME, LAND_AT_SITE, ...), a score (disruption, collision risk 0 to 4) and a noul (is the blockage persistent, is there a traffic conflict). Collisions stay physically possible, so a bad decision costs something.

Five seconds of live decisions in the dashboard, recorded from the screen.
What "reflex" means here. We use the term reflex control for a decision model that answers each agent's situation as it is now: (1) stateless, from the current snapshot only, with no plan and no memory of its last action; (2) fast and per-agent, one answer per drone about once a second (Jev: median 0.56 s); (3) bounded, chosen from a closed set of actions (CONTINUE, HOLD, SLOW_FOR_TRAFFIC, CHANGE_LEVEL and so on), never free-form commands; (4) arbitrated, so deterministic rules can override it and every override is logged. Reflex is contrasted with reasoning, a slower, full LLM thinking step by step, which this experiment does not use. The stack is reasoning (not built), reflex (Jev), and a deterministic backstop beneath it (which runs every 0.1 s tick without asking Jev). The word reflex is already used in robotics for fast low-level control layers; here it is applied to decision models. Jev supervises the mission and does not stabilize flight. A failure specific to this design is the reflex gap: a stateless reflex repeating an unsafe loop (hold, resume, creep), which caused 7 of the 12 collisions in arm A.
What a decision model is. One way to build a decision model is to start with a pretrained language model and replace its text-generation output layer with a decision head: a small neural network that scores the choices supplied with each request. The underlying model still processes the input and its context, but instead of writing an answer word by word, it produces scores for the available options. Training uses examples containing context, possible choices, and the correct or preferred decision; developers train the new head and adjust some or all of the underlying model's parameters. Additional training can improve confidence calibration, so a reported confidence of 80% corresponds more closely to being correct 80% of the time. At inference time, this design can score choices in a single forward pass, avoiding repeated text-generation steps, which can make decisions faster and outputs easier for software to use. This is a reference approach, not a universal architecture, and fast, structured decisions are not automatically safe or correct. In this experiment Jev is used as a black box through its API: we send a state and the admissible choices, and we get scores back. We did not train, inspect or tune the model, and we did not measure its calibration.
What we are trying to show. That Jev's quick answers, with the arbitration, battery and traffic rules listed below, are enough to run the fleet, and, if they are not, how much a deterministic safety layer (the "backstop") has to cover. The backstop only brakes, separates and moves vertically when drones get too close, and every time it overrides Jev the decision is logged as OVERRIDDEN.
The experiment. Same build, same scenario (Max stress: 10 drones, mixed batteries, obstructed legs), same 30 minutes, one run per arm. The only difference is the backstop:

Result in one line. Without the backstop: 12 collisions (24 crashed drones) and 10 missions. With it: 0 collisions and 9 missions. Details below.
- Without the backstop (arm A): 12 collisions, 23 further separation-loss encounters, 353 threat encounters, 10 missions. Jev and the rules resolved 318 of the 353 encounters, but 803 of 6,570 decisions (12 %) were rule overrides of Jev's choice.
- With the backstop (arm B): 0 collisions, 3 separation-loss encounters, 360 threat encounters, 9 missions. The backstop acted 73 times and resolved 43 encounters (12 %).
- Why arm A collided: 7 of 12 collisions were a hold-and-resume creep toward a stationary drone, mostly at launch points. The rule I added (NOUL_YIELD) reads closing speed, which the hold itself sets to zero, and Jev's noul followed the same signal. Another 3 were a priority drone driving into a drone that was holding.
- What this supports: Jev answers fast enough (median 0.56 s, no errors in 13,149 calls) and resolves most encounters, but at this density it needs the safety layer, and no mission cost from the layer is visible in one run.
- What it does not support: a claim that Jev alone is the control brain at 10 drones. It also does not show the backstop is the only thing that would work, since a rule fixing the creep was not tried.
One run per arm, so the intervals are wide and read as indicative. Details, including limits, are in the last section.
Source code. The demo is open source: claudiotx/reflex-drone-flight-control.
Experiment design: two arms, one build
The comparison is two 30-minute runs of the same code on the same scenario (Max stress, 10 drones, loop on). The only difference is the deterministic traffic backstop.
| Arm A: backstop off | Arm B: backstop on | |
|---|---|---|
| Run folder | runs/20261002-185503 | runs/20261002-220525 |
Backstop (BACKSTOP) | 0 | 1 |
| What decides traffic | Jev's choice, plus NOUL_YIELD and the return-to-cruise hold | Jev's choice, those two rules, and the full backstop |
| Scenario / drones / duration | Max stress, 10 drones, 30 min sim time | the same |
| Jev | live, typesafe-jev-1.13.0 | the same |
| Code | the same build; no code edits between the two runs | the same |
| Replicates | 1 | 1 |
The question for arm A is whether Jev's quick choice, score and noul answers can manage traffic and low battery on a shared map without a safety net. Arm B shows what the same fleet does with the net in place, and how often the net acts (counted as OVERRIDDEN rows).
Scenarios in the dashboard. The buttons that select a scenario set the fleet, starting batteries, wind and obstructions. Only Max stress is measured in this report.
| Scenario | What it sets up |
|---|---|
| Mixed fleet | 3 drones, different batteries (100 / 85 / 62 %), wind on; one temporary, one permanent and one noisy obstruction |
| Clear route | 3 drones at 100 %, no wind, no obstruction |
| Temporary obstruction | 3 drones at 100 %, each corridor blocked for 12–22 s |
| Persistent obstruction | 2 of 3 drones have a permanently blocked corridor |
| Ambiguous readings | noisy blocked/clear sensor readings on two drones, a temporary block on the third |
| Jev outage | Jev is made unavailable so the fallback runs; wind on, one permanent and one temporary obstruction |
| Max stress (10 drones) | the 10-drone fleet from the map table: batteries 38–100 %, wind on, 7 obstructed legs |
| Gusty wind | 3 drones, wind zone on, one permanent obstruction |
| Low battery | 3 drones at 52 / 70 / 34 %, wind on, one permanent obstruction |
Controls. Start mission, Reset, Operator hold (turns AI supervision off and holds the drone), Resume AI, Loop (restart missions after each lap and replace crashed drones), and Backstop (on/off, the switch for the two arms).
The map
All 10 drones share one 400 × 300 m map with 11 waypoints, two landing sites, one wind zone and one cruise altitude (45 m). Every drone flies its own 5-waypoint mission through the same waypoints, so routes cross and pile up at shared points.
- Waypoints W1–W11 are shared by every drone. Each waypoint and each landing site is an approved ground landing zone, so a drone can always land somewhere close.
- Home pads sit along the bottom edge (y = 270 m), one per drone, 40 m apart (D1 at x = 30 m, D10 at x = 390 m). A drone parked on a pad or landing site charges at 2 % per second (demo speed).
- Landing sites LS_W and LS_E are extra charging and landing spots at (95, 205) and (305, 205).
- Wind zone is a rectangle from (150, 110) to (260, 200). Inside it the ground speed is capped at 9 m/s and the battery drains twice as fast.
- Corridor obstruction. Each drone has one leg of its mission (leg 2 of 4) that can be blocked. A drone sees the blockage only within 90 m of the leg's entry. Alternates exist: USE_ALT_1 (short detour) and USE_ALT_2 (longer detour, different energy cost).
- Crash recovery. In loop mode a crashed drone is replaced by a new airframe on its home pad after 30 s, so a collision removes two drones for 30 s and the fleet keeps flying.

The map to scale (positions rounded to the nearest character). All 10 drones share the same waypoints and the same 45 m cruise altitude.
The Max stress scenario starts the 10 drones with different batteries and obstructions:
| Drone | Mission | Start battery (%) | Obstructed leg | Obstruction |
|---|---|---|---|---|
| D1 | W1 W2 W6 W3 W8 | 100 | W2 → W6 | clears after 18 s |
| D2 | W8 W3 W7 W4 W2 | 55 | W3 → W7 | permanent |
| D3 | W5 W4 W7 W3 W8 | 70 | W4 → W7 | noisy, clears after 30 s |
| D4 | W10 W2 W3 W4 W11 | 45 | W2 → W3 | clears after 12 s |
| D5 | W11 W4 W3 W2 W10 | 90 | none | none |
| D6 | W2 W8 W4 W7 W9 | 38 | W8 → W4 | permanent |
| D7 | W4 W8 W2 W6 W9 | 80 | W8 → W2 | noisy, clears after 40 s |
| D8 | W1 W8 W3 W9 W6 | 62 | none | none |
| D9 | W5 W8 W3 W9 W7 | 100 | W8 → W3 | clears after 22 s |
| D10 | W3 W8 W2 W10 W6 | 50 | none | none |
The map is a simplified kinematic simulation, not flight stabilisation. All thresholds are demo parameters, not real flight-safety values.
Flight dynamics and rules
Drones fly a simple kinematic model: a cruise speed with limited acceleration, vertical flight levels, and a battery that drains with time and wind. A collision is contact within 4 m horizontally and 4 m vertically. A separation loss is within 15 m horizontally and 8 m vertically at the same moment.
| Quantity | Value |
|---|---|
| Cruise altitude | 45 m (all drones) |
| Cruise speed / max acceleration | 15 m/s / 4 m/s² |
| Climb / descent rate | 4 m/s / 3 m/s |
| Speed inside the wind zone | capped at 9 m/s, battery drains 2× |
| Speed while yielding | capped at 5 m/s (SLOW_FOR_TRAFFIC lasts 6 s) |
| Level change | steps of 10 m, held for 12 s, then back to cruise; minimum level 25 m |
| Traffic shown to Jev | within 150 m horizontally and 35 m vertically, plus any drone climbing or descending toward its level |
| Look-ahead for predicted conflict | 10 s |
| Forced return / forced landing | battery below 25 % / below 10 % |
| Hold limit | 30 s for Jev, 10 s for the fallback |
| Collision / separation loss | 4 m × 4 m / 15 m × 8 m (horizontal × vertical) |
Drone modes are IDLE, TAKEOFF, MISSION, HOLD, RETURN_HOME, DIVERT, LANDING, LANDED, CHARGING and CRASHED. Only drones in MISSION and HOLD ask Jev for traffic decisions. TAKEOFF, LANDING, DIVERT and RETURN_HOME drones are not consulted, but they sweep through every level, and they are visible to other drones as traffic.
Rules that apply in both arms.
- Priority. The drone with the lowest battery has priority (ties go to the lower drone name). A higher-battery drone in a conflict is expected to yield. A drone is never told it has priority over a drone that cannot yield (takeoff, landing, diverting).
- Admissible actions. Jev only chooses among actions that are currently possible. SLOW_FOR_TRAFFIC is offered only if the drone is moving faster than 6 m/s. CHANGE_LEVEL is offered only if a clear level exists.
- Level choice. A candidate level must be at least 8 m from the current altitude, free of other drones in a 200 m radius, and free along the climb path. The level Jev is offered is the one it flies to.
- Return-to-cruise hold. A drone that has finished a level change waits at its level if climbing back would run into traffic within 100 m.
- NOUL_YIELD (added during this work). If Jev chooses CONTINUE or SLOW_FOR_TRAFFIC but its own traffic-conflict noul is high (0.55) or its collision-risk score is high (3.0), and the drone has no priority, a rule changes the action to CHANGE_LEVEL or HOLD. The thresholds are lower for short-contact situations (under 6 s to contact: 0.40 and 2.0) and higher after SLOW_FOR_TRAFFIC (0.70 and 3.5). Every such override is recorded as an OVERRIDDEN decision row.
The backstop (arm B only). A deterministic safety layer that acts when Jev has not resolved a conflict: it brakes a drone so it never closes within 15 m of a same-level drone ahead (from 110 m out), starts a deterministic deconfliction within 50 m, forces an emergency vertical move (10 m/s) when a pair within 25 m is closing, escapes a deadlock after 2 s held by traffic by stepping up a level, and holds a launch until 30 m around the pad is clear.

The flight control engine
Jev makes the decisions; the simulation moves the drones; a thin deterministic layer checks every Jev answer. The simulation ticks every 0.1 s, and each drone in MISSION or HOLD asks Jev about once per second. A call takes about 0.56 s (median), so every decision acts on a state that is already half a second old.
- State. The agent builds a JSON state for the drone: battery, battery needed to return, distance home, next waypoint, corridor observations, altitude and vertical speed, and up to 5 nearby drones in 3D (distance, closing speed, seconds to contact, same level or not, their altitude, vertical speed and target altitude, battery, mode, and whether this drone has priority).
- Admissible actions. Actions that cannot work right now are removed before Jev sees them, so Jev cannot pick a no-op. Each remaining action comes with a one-line consequence (for example, the level CHANGE_LEVEL would fly to).
- Jev answers up to four questions in one call:
maneuveris a choice among CONTINUE, HOLD, USE_ALT_1, USE_ALT_2, SLOW_FOR_TRAFFIC, CHANGE_LEVEL, RETURN_HOME, LAND_AT_SITE and ABSTAIN. This drives the drone.disruptionis a score from 0 to 4 on how much the situation disrupts the mission (blockage, energy, wind).persistent_blockageis a noul (a 0 to 1 degree of belief) on whether a blockage is persistent. Low means clear or clearing.- When traffic is present:
traffic_conflictis a noul that the drone will come within 15 m × 8 m of another within 10 s, andcollision_riskis a score from 0 (none) to 4 (imminent).
- Arbitration. A deterministic check labels each decision ACCEPTED, OVERRIDDEN (a rule replaced the action), REJECTED (not possible) or FALLBACK (Jev failed or was late). Every decision is written to
decisions.jsonlwith the full state, Jev's answers, the verdict, the rule and the action executed. - Execution. The drone applies the action. Answers that arrive after the drone's state has moved on are discarded (epoch check).

The choice drives the flight. Choice probabilities and confidence are recorded but not used. Of the other answers, only the traffic noul and the collision-risk score are used, by the NOUL_YIELD rule. The disruption score and the persistent-blockage noul are recorded and shown in the dashboard but do not change what the drone does.
What the prompt steers Jev toward, by situation (intended use, not measured behaviour).
| Situation | Typical Jev choice |
|---|---|
| Clear route | CONTINUE |
| Corridor blocked, clearing soon | HOLD (hold limit 30 s) |
| Corridor blocked, persistent | USE_ALT_1 or USE_ALT_2, or RETURN_HOME |
| Battery thin for the rest of the mission | LAND_AT_SITE (nearest zone) or RETURN_HOME |
| Crossing traffic | SLOW_FOR_TRAFFIC |
| Head-on or same-line traffic | CHANGE_LEVEL |
| Jev fails, times out or abstains | FALLBACK: the drone holds (hold limit 10 s) |
In arm A (backstop off) the only deterministic actions on traffic are NOUL_YIELD and the return-to-cruise hold. In arm B (backstop on) the full backstop described above also runs. Mock Jev tests only the deterministic floor and was not used for any figure in this report.

Why collisions happened (arm A, backstop off)
Seven of the 12 collisions were the same failure: a drone creeps toward a stationary drone, holding and resuming over and over. I classified each collision from the last six decisions of each drone before contact. These are my readings, not hand-verified for every case.
| Cause | Collisions | What happens |
|---|---|---|
| Hold / continue creep | 7 | The drone holds, so its closing speed drops to 0. Jev's traffic noul falls to about 0.4–0.6 and NOUL_YIELD (which needs closing speed above 1 m/s) releases the drone. It moves, closes again, holds again, and creeps from about 20 m to 4 m. Four involve a drone taking off at a waypoint, two a landing or diverting drone, one a holding drone. |
| Priority drone drives into a holding drone | 3 | The lowest-battery drone keeps going into a higher-battery drone that is holding in its path and cannot get out of the way. Jev rated the conflict high (traffic noul 0.7–0.9) but the priority drone's choice was CONTINUE or SLOW_FOR_TRAFFIC. |
| Late detection at high closing speed | 1 | Head-on at 11–16 m/s. Traffic noul was about 0.2 until 47 m out (about 3 s to contact); the yield came too late. |
| Drone not consulted | 1 | A drone in return-home mode went about 15 s without a decision and was not in the other drone's traffic list. |
The creep is a flaw in the NOUL_YIELD rule I added, not only in Jev: the rule reads the closing speed, and the hold itself sets the closing speed to 0. Jev's own noul follows the same signal. Both 5-drone collisions in the earlier final run (D4×D5 at 661 s and D3×D4 at 1,372 s) were also hold/continue creep at a launch point.

| Time (s) | Pair | Cause |
|---|---|---|
| 381 | D1 × D2 | creep vs takeoff at a waypoint |
| 482 | D1 × D7 | creep vs takeoff |
| 642 | D3 × D7 | creep, both drones |
| 782 | D5 × D8 | priority drone into a holding drone |
| 861 | D6 × D9 | priority drone into a holding drone |
| 1,051 | D6 × D10 | drone not consulted (takeoff vs return-home) |
| 1,081 | D3 × D5 | late detection, head-on |
| 1,186 | D5 × D9 | priority drone into a holding drone |
| 1,187 | D2 × D7 | creep vs landing drone |
| 1,378 | D1 × D5 | creep vs takeoff |
| 1,517 | D1 × D4 | creep vs diverting drone |
| 1,640 | D1 × D8 | creep vs takeoff |
In the logged creep cases each hold-and-resume cycle took about 1 s (one Jev decision) and closed 2–5 m.
30-minute statistics
With the backstop off the fleet had 12 collisions (24 crashed drones) in 30 minutes; with it on, none. Missions were about the same, 10 against 9. Intervals are exact 95 % Poisson intervals on the event counts.
| Measure | Arm A: backstop off | Arm B: backstop on |
|---|---|---|
| Collisions (events) | 12 (95 % CI 6.2–21.0) | 0 (upper bound 3.7) |
| Crashed drones | 24 | 0 |
| Collisions per fleet-hour | 24 (12.4–42.0) | 0 (upper bound 7.4) |
| Separation-loss encounters, no contact | 23 (14.6–34.5) | 3 (0.6–8.8) |
| Separation-loss events, raw count | 38 | 27 |
| Threat encounters | 353 | 360 |
| Resolved without the backstop (Jev choice, plus NOUL_YIELD) | 318 | 314 |
| Resolved by the backstop | not active | 43 |
| Missions completed | 10 (4.8–18.4) | 9 (4.1–17.1) |
| Drones that completed no mission | 6 of 10 | 6 of 10 |
| Jev calls / errors | 6,570 / 0 | 6,579 / 0 |
| Jev latency, median / 95th percentile | 560 / 727 ms | 553 / 738 ms |
| Calls over the 1 s deadline | 17 | 25 |
| Rule overrides: NOUL_YIELD | 803 | 776 |
| Rule overrides: RETURN_TO_CRUISE_HOLD | 54 | 0 |
| Backstop overrides: TRAFFIC_BACKSTOP / EMERGENCY_BATTERY | 0 / 0 | 73 / 1 |
| Rejected: NO_FREE_LEVEL | 58 | 63 |
The collision curve in arm A climbs steadily from minute 7, about one every 2 minutes, so the 12 are not one bad stretch.
How to read the numbers.
- Close calls did not go away with the backstop. Arm B still had 27 separation-loss events against 38 in arm A. The two arms are separate runs whose trajectories diverged, so these counts are not a one-to-one mapping of collisions to near misses.
- 12 % of arm B's encounters needed the backstop (43 of 360). What those encounters would have done without it cannot be read from this run; arm A is the nearest evidence.
- Missions did not change. The two arms are within each other's intervals. Mission output is concentrated: D1 completed 5 missions in arm A and 4 in arm B, and 6 of 10 drones completed none in either arm.
- Both arms are one run. The intervals treat collisions as independent events. They are not (a collision removes two drones for 30 s and the creep failure repeats at launch points), so the real uncertainty is wider than shown. The two runs were three hours apart on identical code.
- Raw separation-loss events and encounters differ. An encounter is one pair interaction and ends in exactly one outcome; the raw count includes repeated loss events inside one encounter.
Limitations and next steps
The result is indicative: one run per arm, one scenario, and rules that were tuned on that scenario.
- One run per arm. The runs were three hours apart. Timing differences (Jev latency of 0.4–0.9 s against a 0.1 s simulation tick) change the trajectories, so a second run of the same arm would differ. Earlier backstop-off runs at 10 drones on older code had 10, 19 (in 23 min) and 15 collisions, which shows the spread; those runs are not comparable and are not used above.
- Jev is stable on the choice but not exactly repeatable. I sent 20 recorded payloads three times each: the choice was identical every time, but noul varied by up to 0.04 and collision risk by up to 0.27. A replay of 100 recorded payloads gave different numbers in all 100 and a different choice in at least 9, for reasons I did not trace (the recorded state may not match what was sent, or the model may have changed). The NOUL_YIELD thresholds (0.55 and 3.0) sit inside that spread.
- Tuned on the test. The rules and fixes were added after watching 10-drone backstop-off runs fail. Arm A is therefore not held-out data. A different scenario would be a cleaner test.
- No rules-only baseline. Arm B shows the full system. It does not show how much of the safety comes from Jev and how much from the backstop alone. A third run with a mock Jev under the same backstop would show that.
- Encounter labels overstate Jev. An encounter resolved after a NOUL_YIELD override is counted as resolved by Jev (318 and 314 above).
- Collision causes are my classification from the last six decisions of each drone, not verified one by one.
- One scenario measured. The other scenarios in the dashboard were not measured for collisions here. The simulation is a simplified kinematic model, not flight dynamics, and its thresholds are demo values.
- One decision model tested. All measured behavior (latency, determinism, jitter, collisions) is Jev's. Claims about decision-model reflex control in general are hypotheses until another model runs the same scenario through the same harness.
- No reasoning arm. The report contrasts reflex with reasoning, but only reflex was run. There is no LLM-reasoning arm, no escalation of mid-confidence cases to a slower model, and no latency benchmark against one. The calibration of Jev's probabilities (NOUL, risk score) was not measured either.
- The backstop uses perfect positions. It reads true drone positions from the simulation. It is not a modeled lidar or any other sensor, so a real backstop would be weaker.
Next steps, in the order I would take them.
- Add a standoff rule for the creep (do not resume toward a stationary drone inside a set distance, and give the holding drone a way out), then rerun arm A.
- Run 2–3 replicates per arm so the intervals reflect run-to-run spread.
- Run the rules-only baseline (mock Jev with the backstop).
- Repeat on a second scenario, such as Mixed fleet or Low battery.
