How iterative, natural-language querying against live fleet, ride, and workforce data — corrected in real time — became the operational plan for a major public event weekend.
A mid-size US city hosts back-to-back public festivals in the same downtown park corridor — one over a weekend, the next running one day longer. The operator managing part of the shared micromobility fleet needed to know: what actually happened last weekend, and what has to change before the next one.
The real value wasn't any single chart — it was the sequence. Each question either built on the last, corrected a wrong assumption, or surfaced an insight that changed the plan. A condensed, anonymized version of that sequence:
The single biggest correction in this whole process was realizing that a data window I'd assumed was a "normal, no-event weekend" was in fact the first festival's actual performance. Every downstream target had been calibrated against the wrong floor.
Before setting any target number, explicitly confirm what conditions produced the comparison data — don't infer it from the calendar.
A worker doing warehouse battery-pack diagnostics produces the exact same signature in the data — high starting charge, no completion reading, negative delta — as a field worker doing genuinely bad swaps. Grading them the same way punishes good workers and hides real problems.
Classify by behavioral pattern (location, start/end values, shift timing) before calculating any performance average across a team.
Neither vehicle model was "better" in the abstract. One consistently earned more per trip in the park/nightlife corridor; the other won decisively downtown. A single fleet-wide revenue-per-ride number would have masked both facts and led to the wrong deployment call.
Always cut performance metrics by the geography or context they occurred in before averaging them into one number.
"The fleet felt low on charge all weekend" isn't actionable. Calculating exact swaps-completed vs. swaps-required against a target average charge turned a vague impression into a specific, schedulable gap — and pointed at exactly which overnight shift needed to change.
Whenever a problem is described qualitatively, look for the underlying quantity that would make the gap unambiguous.
Filtering "rides in the general area" versus rides inside the exact festival boundary polygon, in the exact permitted time window, produced meaningfully different numbers — and the first boundary shape supplied was, on inspection, over a mile from the actual venue.
Verify the geographic or temporal scope of an analysis against ground truth before trusting the output, especially when the source data was hand-entered.
An uncovered overnight window showed up as the same battery-health failure point across multiple event nights in the data. Naming it once wasn't enough — the fix had to be a specific shift reassignment, not a general reminder to "cover nights better."
When a pattern recurs across multiple independent data pulls, treat it as a structural scheduling problem, not a one-off execution miss.
By the time the second festival opened, the plan had shifted from a general "run it like last time" approach to a specific, numbered playbook:
This wasn't a single big report. It was iterative, conversational, and corrected in real time as new data arrived: