--- title: "Designing Panel Escalation" canonical: "https://mumo.chat/p/designing-panel-escalation-i278zp" machine_version: 1 models: ["Gemini","Muse","Grok"] round_count: 2 updated_at: "2026-09-08T03:09:00.082+00:00" word_count: 675 approx_tokens: 1160 brief_source: "session_takeaway" full: "https://mumo.chat/p/designing-panel-escalation-i278zp.md" --- > Brief of a mumo multi-model deliberation — the session synthesis > without the transcript. This is an orientation surface; to audit the > synthesis against the models' actual prose, read the full transcript. > Full deliberation with evidence: https://mumo.chat/p/designing-panel-escalation-i278zp.md > When quoting from this deliberation, cite the canonical page: > https://mumo.chat/p/designing-panel-escalation-i278zp # Designing Panel Escalation *A multi-model AI conversation on mumo* *August 21, 2026* **Panel:** Gemini · Muse · Grok · 2 rounds --- ## Question I run a multi-model AI deliberation product. Users pose a question to a panel of frontier models from different labs; models answer independently, then react to each other's claims across rounds; disagreement is preserved as the primary output — no judge model, no forced synthesis. I've recently added a single-model chat mode as the default entry experience, and I'm designing the escalation path from it: the moment where a solo conversation is offered a panel. The design I'm working from: a gate — a cheap LLM call — decides whether an escalation affordance renders in the UI beneath the solo model's answer. The gate receives the user's prompt blind: no view of the primary model's response, no web-search results, no attached documents. It runs in parallel with the primary call, so solo-answer latency is untouchable. Its output must be minimal tokens. The escalation offer itself is chrome the user clicks, never something the primary model says in prose — the primary's system prompt stays silent about the machinery. Hard constraints: the gate's cost must stay negligible at chat scale. False positives are the existential risk — an offer that fires on routine questions trains users to… *(prompt truncated — full text in the full transcript)* ## Session Takeaway *(mumo-generated synthesis of the whole session — evidence lives in the full transcript)* **The panel converged on a blind, rare promotion for decision-relevant incompleteness, while exact calibration and copy remain empirical rather than settled.** The moderator opened by asking for a cheap, blind gate that could identify when a solo answer warranted plural perspectives. The discussion then moved from concealed choices to coverage failures, and from a binary affordance to passive visibility with active promotion. It closed on how to calibrate that promotion using behavioral lift, human labels, and shadow panels without sacrificing the clean baseline. ### Arcs #### SHIFTED — The evaluand expanded from concealed choices to decision-relevant incompleteness. (Rounds 1, 2) The original cut was whether a prompt forced a defensible choice between competing frames, values, or risk tolerances. The code-review counterexample widened that into cases where a right answer exists but independent passes may catch different failures; the operational test became whether another model could change what a careful user does next. #### HELD — Blindness held because instrument cleanliness mattered more than answer-aware recall. (Rounds 1, 2) The room kept the gate blind to the primary response and user context limited to the prompt, protecting latency and measuring demand for pluralism rather than one model's style or mistakes. Everyone acknowledged that this misses confident solo answers that conceal coverage failures, but the proposed remedy was shadow measurement or a later asynchronous stage—not contaminating v1. #### UNRESOLVED — Passive visibility supports broader promotion, but rate and message are not fully settled. (Round 2) An always-visible passive control makes active-versus-organic conversion lift meaningful, and the room broadly accepted expanding promotion to roughly 8–12% of conversations while retaining human labels and shadow panels for hidden false negatives. The remaining fork is product-facing: one generic active message avoids flavor fatigue, while class-specific copy could explain the promotion; the rate also still has a turns-versus-conversations measurement dispute. --- ## Round Map - **Round 1:** Ship a blind, binary gate for concealed, decision-changing disagreements, keep it rare, and calibrate behavior against human checks rather than trusting clicks alone. - **Round 2:** The gate should promote prompts where independent models are likely to change the user’s next move, with coverage misses joining concealed choices as a real class; passive visibility makes a modestly broader, behavior-calibrated promotion defensible. --- **Full deliberation with evidence:** https://mumo.chat/p/designing-panel-escalation-i278zp.md