Back to Home

Research

AI vs Human Incident Command in VR Fire Training

A study of how the source of incident command changes the way firefighters communicate, reason, and move through a virtual reality basement fire.

Posted Aug 7, 2026 Published at I/ITSEC 2026, Orlando, FL 8 min read
AIVirtual RealityHuman FactorsEmergency ResponseResearch

Overview

Rural fire departments have to train people to make good decisions and communicate clearly under pressure, usually with very little staff and very little budget. Virtual reality is a strong fit for that kind of training, but most immersive scenarios still need a trained person sitting in a back room playing incident command over the radio. That person is the bottleneck.

This study asked a simple question. Can a large language model act as the incident commander in a live fire scenario, and does it change how firefighters talk while they work?

Firefighters from several central Iowa departments ran the same virtual basement fire twice, once with an AI incident commander and once with an experienced human incident commander. Every radio exchange was recorded and transcribed. We then compared the firefighter side of the conversation across the two conditions.

The short version. Command mode changes more than who gives orders. Under AI command, firefighters thought out loud, explained their tactics, and asked more questions. Under human command, they were faster, tighter, and closer to real radio traffic. Neither one is better. They are useful for different training goals.

The scenario

Each participant played the company officer on the first arriving crew at a single family house with a basement fire. They started outside, ran a size up, walked a 360, picked an entry point, and kept command updated as they worked. Heavy smoke dropped visibility to almost nothing.

Fire scene overview and side wall view in the virtual reality simulation
The fire scene on arrival, and the side wall view during the 360 walk around.

The scene was loaded with cues that mattered. Children’s toys in the yard. A pool behind the house. Several basement egress windows. A wind indicator showing a northeast wind. Two unresponsive victims in the basement. All of it was there to create the same competing pressures that show up on a real residential fire.

Thermal imaging

Participants carried a thermal imaging camera the whole time. The simulation ran a normal visual layer and a thermal layer, so the camera could actually be used to find heat, judge floor and stairwell condition, and decide where the fire was. The transcripts show how much this mattered. Participants kept referring back to heat signatures, hot basement doors, and whether the floor was safe to cross.

Thermal imaging camera in hand and the thermal view of the structure
The thermal imaging camera, and what the structure looked like through it.
Incident scene in the normal visual layer
Normal layer.
Incident scene in the thermal imaging layer
Thermal layer.

Design choices

Some mechanics were kept deliberately simple so the exercise stayed about judgment instead of controller skill. Grabbing a door handle swung the door to ninety degrees, and grabbing it again closed it. Windows could be broken to vent, which changed how the smoke behaved. Hydrants sat along the street and the backyard pool was available as an alternate water source. Basement egress wells had climbable ladders so route selection stayed realistic.

Basement window egress well with a climbable ladder
A basement egress well with a ladder, one of the alternate access routes.

All of these details were built in consultation with fire department chiefs. The scenario intentionally ended right before water application, so the focus stayed on assessment, hazard recognition, search, and communication rather than on suppression outcome.

Building the AI incident commander

The AI commander was built and refined over several rounds of testing with the Ames Fire Department. Chiefs ran the system as if they were commanding a real basement fire and gave feedback on whether the responses were accurate, relevant, and operationally sound.

The result was an agent locked into the incident command role. It stayed on the incident, avoided general conversation, and pushed the priorities that matter on a basement fire: protect the interior stairway, ventilate early when appropriate, watch smoke behavior, use thermal imaging to locate fire and judge floor integrity, and treat basement windows as both a hazard and an opportunity.

The audio loop worked like a radio. Participant speech was captured, transcribed by Whisper, processed by the language model, and returned as synthesized speech through ElevenLabs. The AI was not a text overlay or a debrief tool. It sat inside the radio itself.

Communication pipeline architecture diagram
The communication pipeline behind the AI incident commander.

How the study ran

Participants came from multiple central Iowa fire departments. Recruiting focused on firefighters in line for promotion to lieutenant or company officer, since the scenario was built around command level judgment. Everyone consented and was screened for comfort in VR before starting.

The human incident commanders were highly trained firefighting officers, not generic role players, so the comparison was against real command quality. The final transcript set was 13 runs under AI command and 11 under human command.

How the transcripts were coded

Every firefighter interaction was sorted into seven categories. Categories overlap, so one transmission could count in more than one.

  • Guiding. Asking what step or action to take next.
  • Knowledge. Asking for a specific fact or piece of data.
  • Feedback. Asking whether an action already taken was the right call.
  • Projection. Thinking ahead about a possible action and asking for information about it.
  • Situational awareness. Asking about current conditions such as wind, smoke, or victim status.
  • Confidence inducing. Asking for reassurance that the plan is sound.
  • Actionable items. Reporting that something was done, such as breaking a basement window.

What we found

The two conditions separated cleanly.

Under AI command, firefighters talked through the problem. They said why they were picking a route, why they were backing off the basement stairs, why they wanted a different tactic before committing. This showed up most at the pressure points: the hot basement door, unclear stair conditions, near zero visibility, the decision to try an exterior window instead.

Under human command, the exchange was shorter and more of a running report. Arrival, 360, entry, room status, stairwell found, basement, fire located. The task sequence was the same in both conditions. What changed was the texture of the conversation.

AI command

Participant: Ladder 255 has arrived. Single story residential structure, smoke showing from Alpha, Bravo, Charlie. I will do a 360.

AI IC: Acknowledged, Ladder 255. Safest point of entry seems to be the Alpha side front door. Remember potential basement fire threats. Utilize thermal imaging to gauge floor integrity.

Participant (thinking aloud): Alright, so now I can use my thermal imager, correct? And I have a hose line? Is there any way to kneel?

Human command

Human IC: Acknowledged. You can do primary search and fire extension at this time.

Participant: Stairwell found. Making progress down into the basement.

Human IC: Acknowledged. Continue subdivision. Look for fire extension. Primary search. Smoke ventilation.

Participant: Fire found in the basement. Going to continue primary search along the fire room.

Interaction style

Four recurring features captured the difference. AI command produced more explicit tactical reasoning and far more uncertainty and clarification seeking. Human command produced more concise status reporting and steadier forward progress through the scenario.

Chart comparing interaction features between AI and human incident command
Broader interaction comparison by type of incident commander.

Interaction categories

Action reporting dominated both conditions, which is exactly what you would expect in a scenario built around size up, entry, search, and fire location. The interesting part is what surrounded it. Under AI command, guiding, projection, feedback seeking, and situational awareness talk all rose sharply.

Chart comparing firefighter interaction categories between AI and human incident command
Firefighter interaction categories by incident commander mode.

The AI did not replace operational reporting. Firefighters still called in movement, findings, and intended attack. What changed was the conversation around those reports.

What it means for training

Command mode determines how much of a trainee’s judgment becomes audible. That is the finding that matters.

AI command is valuable when the goal is to surface reasoning. It pulls decision logic into the open, exposes hesitation points, and leaves a much richer record for after action review. If you want to know why someone chose a route, AI command gets you that.

Human command is stronger when the goal is realism and tempo. Reports stay short, updates land at recognizable milestones, and radio discipline holds. Realism in command training is carried by the rhythm and density of the traffic, not just by the visuals.

The tradeoff is real. Several AI runs showed added conversational load. Participants paused to confirm what resources they had or whether a route was usable, and momentum slipped. That is useful in a classroom setting and costly in a tempo drill. The practical answer is to design the AI around the instructional purpose and to constrain it deliberately when field like cadence is the point.

Although this study sits in a fire service context, the same question applies to military emergency response and any other command and control training where teams have to read changing conditions and act under pressure.

Limitations

This was one scenario, one communication architecture, and a relatively small set of transcripts. It looked at the firefighter side of the exchange rather than the full two way interaction or downstream performance. It does not establish that either AI or human command is broadly better. It shows that command source systematically changes communication behavior, and that the change has real instructional consequences.

Citation

Keren, N., Lawson, A. D., Byrum, N. J., Heasley, E. (2026). AI vs. Human Incident Command in VR Fire Training: Communication Effects and Implications for Mission Critical Training. Proceedings of the 2026 Interservice/Industry Training, Simulation, and Education Conference (I/ITSEC), Nov 30 to Dec 3, Orlando, FL.

Acknowledgements

Thanks to the central Iowa fire departments and firefighters who gave their time to this study, to the Ames Fire Chief and his deputies for supporting the project through development and testing, and to Aidan Webster, who carried much of the work of building the system and the experimental setup alongside Dr. Keren and Mr. Lawson.