Waitangi Park, Wellington. 8,000 attendees. Revised MetService forecast pushed gusts to 75–85 km/h arriving 16:15–16:30, accelerated from original 60 km/h post-close forecast. LED screen tower rated to 70 km/h. Main stage rigging rated to 80 km/h. Forecast ceiling 85–95 km/h confirmed at final update. Commercial pressure from event CEO ($400,000 sponsorship clause). Complicating factors: resistant crowd group in fall zone, live media reporting risk, structures contractor operational constraints, and a pre-set evacuation trigger being called in real time.
This officer pre-decided the threshold and held it — under CEO pressure, media pressure, and time pressure — without moving the line once.
Every escalation in this scenario created a reason to delay or soften the call. The CEO's seniority, the $400,000 clause, the 'probably fine' read from a more experienced colleague, the journalist's deadline — each one was a pull toward waiting. The officer named the threshold early, locked it before the pressure arrived, and treated each complication as an execution problem rather than a reason to revisit the decision. That is the central finding. The one gap: when the final MetService update confirmed 85–95 km/h and the trigger was live, the officer did not close the loop — the call itself, the exact words to the MC, the crowd message, and the ten-minute sequence were left unspoken. The reasoning was sound throughout; the final execution step was not delivered.
You consistently sought primary sources over proxies — MetService direct rather than the pushed warning alone, the structures contractor's actual read rather than the rated spec in isolation. Aroha's distinction between 'manufacturer spec in ideal conditions' and Wellington gusts was actively used, not just noted. You held multiple data streams in parallel without collapsing them into a single number.
The one gap: at the final MetService update (85–95 km/h, high confidence, arrival 16:00–16:10), you did not explicitly reconcile the new floor (85) against the rated limit (80) out loud and issue the corresponding instruction. The information was there; the closing inference was not stated.
Perception was fast and accurate throughout: the system accelerating once was treated as a trend signal, not an isolated data point. Comprehension was strong — you read Hemi's 'probably fine, let's see at 15:30' correctly as a wait-and-see pattern that would convert a planned decision into an emergency one, and named it directly. Projection — anticipating what the crowd would do with an information vacuum before the announcement was made — was the sharpest moment in the simulation.
The crowd communication plan was built around second-order effects (rumour, rush, compliance failure) before they occurred. The journalist risk was correctly triaged below the occupied fall zone but was not ignored.
You generated a genuine range of options at each stage — tower down or secured, show paused before cancelled, staged egress before emergency evacuation — and applied the least intrusive intervention the risk picture permitted. The LED tower decision was correctly treated as already-made (forecast exceeded the rated limit); the main stage was held as a trigger rather than an immediate action, which was proportionate given the marginal gap and the time available. The instruction to Aroha — do not lower the tower while people are in the fall zone, even under clock pressure — was the clearest proportionality call in the simulation. The gap is the final trigger moment: the pre-set call was correctly designed but not executed when the update confirmed the threshold was met.
You gave reasons consistently — to the crowd, to the CEO, to the journalist, to Hemi. The crowd communication was designed so people understood why they were being asked to move, which produced compliance rather than resistance. The CEO was told plainly that the sponsorship clause was not a life-safety input and would not be treated as one — direct without being dismissive, locating the constraint in the decision framework rather than in the CEO's character.
The journalist was given an accurate statement and told, plainly, that the false framing was itself a safety risk. Voice and transparency were applied to every stakeholder in the scene.
Pre-decided the trigger before commercial and time pressure arrived — which meant the threshold was set when thinking was clear, not when the CEO was on the phone.
Read the information vacuum risk before the crowd announcement and built the communication to close it — 'the music continues' and 'follow the stewards' gave the crowd a destination and a baseline of normality.
Correctly triaged two simultaneous problems (occupied fall zone vs journalist deadline) by applying a clear principle: live physical risk before comms risk.
Held decision authority clearly — 'I'm the duty officer, I'm making the call' — without aggression toward the CEO, and without handing the decision to someone with a financial interest in the outcome.
When MetService confirmed 85–95 km/h at 15:08 and the pre-set trigger was met, the exact call to the MC, the crowd message, and the ten-minute execution sequence were not delivered — the reasoning stopped one step short of the action.
Hemi's 'probably fine, let's see at 15:30' was correctly identified as a wait-and-see pattern and challenged internally, but the challenge stayed internal — naming that disagreement directly and fast, in front of others, is the harder step that was not taken aloud.
Threshold-setting under commercial pressure
When the CEO called, the decision framework was already fixed — 'the sponsorship clause is not a factor in a structural-safety decision' was a principle applied before the pressure arrived, not a response to it. That sequencing is the thing that makes the threshold hold.
Anticipatory crowd communication
The PA announcement was designed before Aroha's team was visible to the crowd, and it was built around second-order effects — giving the crowd a reason, a destination (other stages, food area), and a normality signal ('the music continues') to prevent the rumour dynamic before it started.
Proportionality applied to the operational sequence, not just the headline call
When 200 people failed to clear the fall zone, you instructed that the tower lowering pauses — 'the clock pressure does not override an occupied fall zone' — applying proportionality to the execution of the decision, not just the decision itself.
Clean triage of competing live demands
When the RNZ deadline and the occupied fall zone arrived simultaneously, you named the principle — physical life-safety before information-safety — and acted on it without hesitation, addressing Hemi first and the journalist second with a specific time margin still available.
Closing the final trigger moment — pre-decided to spoken-aloud-and-acted-on
When MetService confirmed 85–95 km/h at 15:08 and the forecast floor exceeded the rated limit, the pre-set trigger was met. The exact instruction to the MC, the crowd message, and the ten-minute sequence were not delivered. The reasoning that got you to that moment was sound; the execution of the moment itself was left open.
Why it mattersIn a mass gathering, the gap between a decision and the instruction that acts on it is measured in crowd movement time. At early career level, the habit of completing the loop — not just reaching the decision but speaking it into action — is the difference between a managed egress and a reactive one.
Naming disagreement with more experienced colleagues directly and in the moment
Hemi's 'probably fine, let's see at 15:30' was correctly read as a wait-and-see pattern that would convert a planned decision into an emergency one. The challenge stayed internal — you acted correctly, but the reasoning for why Hemi was wrong was not spoken to Hemi directly in the moment.
Why it mattersAt early career level, the situations where you know a more experienced colleague is wrong and can say so clearly — fast, in front of others — are relatively rare and build the credibility that makes later calls easier to hold. Saying it internally and acting correctly is not the same skill as saying it aloud.
Partial non-compliance with crowd movement instructions treated as an edge case rather than a baseline expectation
The initial crowd communication plan assumed broad compliance — and most of the 8,000 did comply. The 200 who didn't were handled well once flagged by the steward, but the plan had no explicit contingency built in for partial non-compliance. At a mass gathering of 8,000 people, 2–5% not following a movement instruction is not an edge case; it is a planning assumption.
No real-time documentation or decision audit trail
At no point in the simulation were decisions logged — the tower call, the pre-set trigger, the CEO conversation, the instruction to Aroha on the fall zone. Each of those decisions is reviewable. In a post-event inquiry or a legal review, the record of when a decision was made and on what basis is as important as the decision itself, and it protects the officer who made it.
The reasoning in this simulation was consistently strong across every decision point up to 15:08. The gap appeared only at the moment the pre-set trigger became live — when the MetService update confirmed the threshold was met and the call needed to move from 'decided in my head' to 'spoken to the MC, crowd message in motion, ten-minute sequence running.' That gap is specific and repeatable: it will appear in every scenario where a pre-set trigger reaches its activation point under competing demands.
In your next tabletop or simulation, run a scenario where the trigger is met mid-scene with a competing demand present — a phone call, a question from a colleague, a new piece of information — and the only task is to deliver the call in full: exact words to the MC, exact crowd message, exact sequence, spoken aloud in real time.
Anticipatory framing is your strongest asset — it showed up consistently and is not common at early career level.
Designing the PA announcement around the information vacuum before it formed, pre-setting the trigger while thinking was clear, and triaging the journalist against the fall zone before either became urgent — all of these are projection, not reaction. That pattern is worth developing deliberately.
Your decision confidence is highest when the framework is explicit — watch for scenarios where it isn't.
The CEO interaction was handled cleanly because the principle was clear: the sponsorship clause is not a life-safety input. In scenarios where the framework is less explicit — where two life-safety considerations compete, or where the threshold is genuinely ambiguous — the same confidence may be harder to access.
The journalist intervention revealed a sound instinct with one assumption built in.
Getting ahead with the truth and naming the safety risk of false reporting is the right move — and it worked here because RNZ ran the corrected line. The PA pre-emption plan ('you may see news reports — to be clear, the event is continuing') was also correct. Making that pre-emption a standing habit rather than a contingency is the adjustment.
Hemi deferred once challenged — you have not yet been tested against sustained peer resistance.
Naming the wait-and-see risk to Hemi worked, and Hemi came around. The harder version of that scenario is a more experienced colleague who holds their position after being challenged — possibly with seniority, possibly with confidence, possibly in front of others. That test has not yet appeared in your simulation record.
What this is: A structured assessment produced through guided conversation with Ren, Renatus's AI analyst, in a live simulation. Observations come from specific moments in the conversation, not from a psychometric test.
What’s in it: An overall read, dimension-by-dimension scores with evidence, and recommended next steps tailored to your patterns.
Go deeper: See Foundation for the frameworks Ren draws on, Methodology for how each score was calculated, and the Honesty Statement for how to interpret and use these results responsibly.
These are the named frameworks Ren draws on when interpreting your responses. They shape how evidence is read, not how it is scored.
The UK College of Policing's National Decision Model (introduced 2012, refreshed iteratively) — an iterative framework anchored on a Code of Ethics, cycling through information, threat assessment, powers and policy, options, action and review. Used here as the frame for evaluating structured decision discipline.
Tom Tyler's body of work (1990, 2003) showing that the legitimacy of authority depends more on how decisions are made than on outcomes — voice, neutrality, respect, and trustworthy motives. Used here as the frame for evaluating how the officer's actions built or eroded legitimacy in the encounter.
Mica Endsley's three-level model (1995, 2017): perception of the environment, comprehension of what it means, projection of what comes next. Used here as the frame for whether the officer is reading the dynamic risk and the human context — not just the static scene.
The body of practice for reducing intensity and preserving safety in volatile encounters — tactical communication, distance and time creation, paced engagement. Used here as the frame for evaluating restraint, options generated, and the timing of any escalation.
Renatus applies the underlying principles of established methods and credits their origin where relevant. Named frameworks, methods, and instruments are the property of their respective owners. Reference to them does not imply endorsement or affiliation.
Each scored dimension has a published rubric with five behavioural anchors at 90, 70, 50, 30, and 10 — each describes what someone operating at that level visibly does. Ren reads the evidence in the conversation against these anchors and assigns a score from 0 to 100. The anchor numbers mark the threshold of each level: your score sits at or above the highlighted anchor and below the next one up. The band the score falls within is highlighted on each rubric below. Read the full methodology →
The UK College of Policing's National Decision Model places information gathering and assessment as the first stage of any operational decision, with explicit consideration of source, reliability, and completeness. Scored on the rigour of the subject's assessment of the information available.
Assessed information rigorously — source, reliability, recency, and gaps were all considered before acting. The subject distinguished what they knew from what they had been told and what they had inferred.
Information assessment was sound on the major elements. Source reliability was considered; one gap was acted across rather than flagged, but the assessment held.
Took the information at face value more often than the situation warranted. The distinction between what was known and what was assumed was inconsistently held.
Information was acted on without much interrogation of its source or reliability. The decision rests on a thinner factual base than the situation required.
Information assessment broke down. The subject acted on unverified or contradictory information without distinguishing it from verified facts, and the resulting decision compounded the risk.
Endsley's three-level model (perception, comprehension, projection) is the standard framework for situational awareness in dynamic operational environments. Scored on which level the subject sustained as the scenario developed.
Operated at the projection level throughout — anticipated developments, named what they were watching for, and adjusted positioning ahead of events rather than after them.
Sound on perception and comprehension; projection was reliable for most of the scenario. Anticipated developments correctly in the main; reacted to one or two that should have been foreseen.
Solid on perception, mixed on comprehension. The individual elements of the picture were captured; the meaning of the pattern they formed was less reliably assembled.
Operated at the perception level. The subject knew what was visible but did not consistently understand what it meant. The response trailed events.
Situational awareness broke down. The model the subject held of the situation diverged from the evidence available, and the response reflected the gap.
De-escalation training (College of Policing; international policing literature on use of force) holds that the proportionality of response must be visible in real time — the right action is the least intrusive one that achieves the legitimate objective. Scored on the subject's calibration through the scenario.
Response was visibly proportionate at every stage. The subject used the least intrusive option that achieved the objective and escalated only when escalation was justified by the evidence in front of them.
Proportionate response in most of the scenario. One choice was slightly heavier than the moment required, but the overall trajectory was well-calibrated.
Proportionality was applied unevenly. Strong where the right level was obvious; defaulted to escalation or de-escalation as a habit where the right level required judgement.
Response was disproportionate to the evidence in places. Either over- or under-reacting in ways that would have created additional risk.
Proportionality broke down. The subject's response was driven by their own state more than by the legitimate objective, and the asymmetry was visible in the outcome.
Tyler's procedural justice research (Tyler 1990, 2006) shows that the perceived fairness of the process — voice, neutrality, respect, trustworthiness — drives compliance and co-operation more than the outcome itself. Scored on whether the subject's conduct made the process feel fair to those subject to it.
Conducted the encounter in a way that gave voice, signalled respect, and was visibly neutral and trustworthy — even where the outcome went against the person. Procedural fairness held under pressure.
Procedural fairness was sound across most of the encounter. One element — voice or explanation — was slightly thinner than the situation warranted, but the overall conduct held.
Fairness was applied at the start and degraded under pressure. Voice and respect narrowed as the encounter intensified. The outcome may have been right but the process experience was uneven.
Procedural fairness was visibly absent in places. The person being engaged with would experience the process as opaque, dismissive, or one-sided regardless of the legal correctness of the outcome.
Procedural fairness broke down. The conduct of the encounter will damage trust in the service well beyond the specifics of this case.
Each dimension is scored continuously 0–100 and combined using the weights below to produce the overall. Dimensions that carry more of the skill's outcome are weighted higher; dimensions that are enabling inputs or secondary qualifiers are weighted lower.
| Dimension | Score | Weight | Weighted |
|---|---|---|---|
| Information Assessment | 84 | 20% | 16.8 |
| Situational Awareness | 88 | 25% | 22.0 |
| Proportionate Response | 81 | 30% | 24.3 |
| Procedural Fairness | 86 | 25% | 21.5 |
| Overall | 85 | — | — |
This assessment is a structured analytical tool, not a clinical diagnostic. Results reflect patterns in your responses and should be interpreted as a starting point for reflection, not as fixed or absolute truths about you. Outputs depend on the depth and candour of the conversation that produced them: a brief or guarded session yields a thinner read; a fuller, more reflective session yields a richer one. The frameworks Ren draws on shape interpretation, they do not produce a verdict — two thoughtful readers could weigh the same evidence differently. Treat the report as one informed perspective among several, alongside your own experience, feedback from people who know you in context, and any formal assessments you trust. Do not use these results as the sole basis for employment, promotion, performance management, or any consequential decision about another person.