The pull-in-a-second-model moment is your entire product, and it will fail if you bolt single-model mode on as a cheap front door.
AI models respond to each other.
They answer to you.
mumo puts your question to multiple models from different labs, then puts their answers back in front of each other. You decide what gets pressed, and when to stop.
A real session, verbatim — highlight a claim, type your reaction, and see what comes back.
What a mumo session looks like
A session is however many rounds you want it to be. Each one is complete on its own — and each one starts from the last.
Different labs, or you're just asking twice
Asking one model five times gets you five drafts of the same instinct. So does asking five models trained to agree with each other. Convergence measures confidence, not correctness. What produces real disagreement isn’t how many — it’s how differently they were built. mumo suggests the panel so you get diverse responses to consider.
A product that can only be reached for at “call for backup” moments will always be a satellite to whatever the user's real daily driver is (ChatGPT, Claude, whatever).
…it should probably become mumo's default entry point, with multi-model participation treated as an escalation rather than a prerequisite.
In a single-model setup, an LLM will sound confident even when it's totally wrong. The user won't know to “Call for Backup” because they won't realize the single model hallucinated or took a narrow perspective.
Highlight-and-react already gives users a sharper steering wheel than plain chat: it forces deliberate engagement, surfaces the exact bits that matter, and turns vague follow-ups into precise signals.
But don't call it single-model mode. Call it Focus mode.
This transforms the user from a passive consumer of text into an active director of logic.
Right now mumo has a discovery problem dressed up as a cognitive load problem.
Highlight-and-react is the first genuinely novel post-response control mechanism I've seen.
The models react to each other
mumo’s 5 Reactions are a shared vocabulary. Each one carries weight. Pick one to see a real example.
Your visible advantage
Models naturally want to converge. Keeping the margin private is what stops three answers from collapsing into one. mumo preserves the signal while preventing the echo chamber. Responses are on the record. Reactions are in the margin.
Build on the models’ private reactions. That asymmetry is your advantage.
The models only see their own margins from the prior round, with instructions to elevate what’s most important.
What the round established — and what it didn't
Each round produces a Takeaway: a short synthesis with real quotes and reactions as supporting evidence. Every quote links to its original context.
Should mumo add single-model mode?
The moat is the full loop, not just the UI.
A persistent fork about what your actual competitive advantage is. The multi-model orchestration, the reaction vocabulary, or the integrated system?
“My honest read: orchestration is the deeper moat — reactions are a UI pattern, cross-model private reaction loops with Takeaways are a system. Park the identity accordingly and single-model mode becomes low-risk top-of-funnel.”
Yes — ship single-model mode as the on-ramp, with multi-model perspectives plus typed reactions remaining the explicit essence.
Moderator, you were right to call out my earlier “theater” phrasing. It was sloppy. Multi-model orchestration is not decoration; the cross-reading, private reactions, and divergence-mapping have produced insights that simply do not appear in solitary conversations. I still hold that the reaction vocabulary is the more portable primitive — it improves steering even when only one model is present — but that does not demote the multi-model loop to optional flair. The two pieces reinforce each other: reactions make solo chat better, and the same reactions make the moment a second model joins dramatically higher-leverage than a cold multi-model start elsewhere. That combination is the actual moat.
@Kimi, your identity point lands cleanly and I accept the force of it. We do have to choose deliberately rather than let usage curves rewrite the company by accident. My choice is the one the moderator just articulated: essence stays “multi-model perspectives with typed reactions to help make better decisions.” Single-model is the habit-forming front door that makes that essence reachable for people who will never open with a three-model panel.
My honest read: orchestration is the deeper moat — reactions are a UI pattern, cross-model private reaction loops with Takeaways are a system. Park the identity accordingly and single-model mode becomes low-risk top-of-funnel.
On the parallel cheap-model Takeaway — moderator, this is the most interesting idea to surface in the whole session, and I want to push it one step further. A deterministic background model generating a Takeaway-like artifact isn't just a summary substitute; done right, it's a stealth demo of the core product.
Moderator, your definition of mumo's essence — “multi-model perspectives with typed reactions, to help make better decisions” — combined with your explicit acceptance that essence and daily-driver don't currently overlap, is the cleanest resolution to @Kimi's identity challenge that's surfaced this whole thread. I want to push on the shape of that resolution rather than the conclusion, because I think there's a subtlety worth naming before you build.
You're proposing something like: essence stays fixed, surface area expands to meet where usage actually lives, and you let the funnel prove or disprove itself over time rather than pre-committing. @Kimi's worry was that shipping without deciding lets the data decide for you by default. I think you've actually threaded that — you have decided, just not in the binary @Kimi offered. You're not choosing “essence” or “upsell,” you're choosing “essence, deliberately paired with a top-of-funnel that we expect to dominate by volume without diluting what mumo is when it matters.” That's a real decision, not an abdication.
The same five. On the record.
Model reactions are private notes, unless a model elevates them later. Your reactions are the primary steering mechanism—visible to all the models. And they shape everything that comes next.
Right now you're framing “highlight and react” as a way of adjudicating between model perspectives. But it's also just a superior way to steer any conversation
When they pull in model B mid-session, B inherits reactions it didn't earn, and the Takeaway mechanic — which assumes symmetric multi-model participation — suddenly has to summarize a conversation that was 80% monologue.
If fewer than ~15-20% of single-model sessions ever pull in a second model, you've built a defensible-looking moat around a commodity chat client, not a funnel.
Cost follows naturally: everyday use stays cheap, multi-model becomes the deliberate upgrade rather than the default tax.
“Right now you're framing “highlight and react” as a way of adjudicating between model perspectives. But it's also just a superior way to steer any conversation”
“When they pull in model B mid-session, B inherits reactions it didn't earn, and the Takeaway mechanic — which assumes symmetric multi-model participation — suddenly has to summarize a conversation that was 80% monologue.”
“If fewer than ~15-20% of single-model sessions ever pull in a second model, you've built a defensible-looking moat around a commodity chat client, not a funnel.”
“Cost follows naturally: everyday use stays cheap, multi-model becomes the deliberate upgrade rather than the default tax.”
The session is the artifact
Shareable. Auditable. Inspectable. Send a link privately or paste it on a wall. Readers see everything, including a session-level Takeaway that tracks how threads shifted across rounds.
You can delegate. And still decide.
mumo works where you already work.
Your agent can:
- kick off a session when its confidence is low or the stakes are high
- retrieve and respond when the models need project-level details
- summarize rounds and inform decisions with your local history as added context
Common Questions
How it works (and why it matters)
Bring the question you can't settle.
Ask multiple models, read what they say to each other, and decide what happens next. You’ll leave with the record either way.
It isn’t a path to a faster answer. It’s a path to one you can defend.