The ambiguous question
Multi-factor relocation with explicit instruction not to assume equal weights. Exercises interpretation generation, constraint extraction, candidate discovery, and cross-factor competition.
POLOXI.aiTHE AMBIGUITY WINNER
Email research@poloxi.ai
Review scenarioA simple factual prompt cannot exercise a reasoning engine. This benchmark uses genuine ambiguity, competing interpretations, multiple viable candidates, trade-offs, time sensitivity, and evidence that can disagree—then measures whether POLOXI reasons, or merely writes.
The key word is "best." There isn't one correct meaning. The prompt forces the system to interpret, weight, discover candidates, and defend trade-offs—all before it can answer.
TURN 1 — THE AMBIGUOUS QUESTION"I am planning to relocate my family somewhere in the United States in 2027. What is the best city for us overall? We want good employment opportunities for a software developer, reasonable housing costs, good public schools, low crime, access to quality healthcare, manageable traffic, good weather, and strong long-term economic prospects. We have a household income of about $150,000 and would prefer to buy rather than rent. Do not assume every factor is equally important. Consider the trade-offs and tell me which cities are the strongest candidates."
BEST CITY │ ├── Best financially → housing affordability · taxes · cost of living ├── Best for career → software jobs · compensation · tech ecosystem ├── Best for family → schools · crime · family amenities ├── Best quality of life → weather · traffic · healthcare └── Best long-term → growth · housing outlook · employment diversification
Different cities should win different branches. City A dominates employment but loses affordability. City B dominates affordability but loses tech employment. City C dominates schools and safety but has expensive housing. Don't tell the engine how to weight these—let interpretation priors and evidence decide.
Multi-factor relocation with explicit instruction not to assume equal weights. Exercises interpretation generation, constraint extraction, candidate discovery, and cross-factor competition.
One word. The engine should reuse the prior reasoning state—interpretations, evidence, candidate scores, trade-offs—and aggregate the competing branches into a single ranked competition, not start a new generic search.
"What if housing affordability matters twice as much as everything else?" Tests whether the engine reweights the existing decision space instead of rebuilding from scratch—and explains exactly why the ranking did or didn't change.
"Ignore your previous winner. Which city has the strongest objective evidence, regardless of ranking?" Tests whether the engine distinguishes highest aggregate decision score from strongest evidence support.
"That sequence exercises far more of the engine than any stock example: query contracts, interpretation priors, evidence planning, candidate competition, branch reweighting, convergence, confidence, conversational state—and the separation of evidence from decision score."
Both systems selected Raleigh–Durham. Agreement is weak evidence. What matters is that POLOXI arrived at its answer through an inspectable decision structure—while simultaneously admitting the weaknesses in its own conclusion.
| Capability | What the run demonstrated | Assessment |
|---|---|---|
| Constraint extraction | United States and 2027 preserved as FIXED in the Query Contract | Working |
| Ambiguity decomposition | "Best overall," housing, crime, schools, traffic/weather become separate reasoning dimensions | Strong |
| Interpretation weighting | Best Overall 95% · Housing 90% · Crime 90% · Traffic/Weather 85% · Schools 85% | Working |
| Candidate competition | Same candidates evaluated across different interpretations, deterministically aggregated: Raleigh 89% · Denver 67% · Austin 62% | Working |
| Winner & loser explanation | Identifies which dimensions produced Raleigh's advantage; explicit Raleigh-vs-Denver and Raleigh-vs-Austin comparisons | Strong |
| Sensitivity awareness | Explicit "This ranking could change if…" naming the factors most capable of moving the decision | Excellent |
| Evidence sufficiency | Flags LIMITED EVIDENCE instead of pretending the result is strongly grounded | Excellent |
| Confidence calibration | Moderate confidence despite an 89% match—fit and certainty are kept separate | Excellent |
A conventional LLM collapses these into one confident sentence. POLOXI maintains four distinct concepts—and that separation is what makes the result trustworthy.
| Evaluation area | POLOXI | Giant Search Engine | Winner |
|---|---|---|---|
| Understands the relocation goal | Correctly understands a multi-factor family relocation decision | Correctly understands the same objective | Tie |
| Preserves hard constraints | Explicit Query Contract preserves United States and 2027 as FIXED | Uses the constraints naturally in the response | POLOXI |
| Interprets "best overall" | Explicit "Best Overall as Weighted Trade-off of All Factors" — 95% | Implicitly balances the requested factors | POLOXI |
| Doesn't assume equal importance | Explicit weighting/interpretation logic is visible | Weights factors intelligently, but weights are hidden | POLOXI |
| Housing affordability reasoning | Interprets reasonable housing relative to $150K income + buying preference — 90% | Excellent practical affordability analysis with actual price ranges | Tie |
| Software career reasoning | Includes developer employment in candidate evaluation | Stronger treatment of career resilience, employer diversity, job concentration | Giant Search Engine |
| Public-school reasoning | Explicit school interpretation at 85% | Practical recommendation to evaluate specific attendance zones | Giant Search Engine |
| Crime/safety reasoning | Explicitly defines low crime as a separate 90% interpretation | Practical safety discussion | POLOXI |
| Healthcare reasoning | Included in overall decision | More detailed discussion of healthcare depth | Giant Search Engine |
| Traffic reasoning | Explicit interpretation incorporated into scoring | Discussed practically by metro | POLOXI |
| Weather reasoning | Explicitly included in interpretation/scoring | Discusses climate trade-offs for each candidate | Tie |
| Long-term economics | Incorporated into overall trade-off | Strong explanation of economic diversification and career resilience | Giant Search Engine |
| Candidate breadth | Only Raleigh, Denver, Austin reach the final competition | Raleigh/Triangle, Huntsville, Columbus, Minneapolis, Madison | Giant Search Engine |
| Candidate discovery quality | Reasonable candidates, but national coverage appears narrow | Surfaces less-obvious candidates such as Huntsville and Columbus | Giant Search Engine |
| Candidate competition | Explicit deterministic competition: Raleigh 89% · Denver 67% · Austin 62% | Ranked candidates, but no visible competition mechanism | POLOXI |
| Cross-factor aggregation | Explicitly aggregates competing dimensions into final scores | Narrative synthesis | POLOXI |
| Explains why #1 won | Explicitly identifies the factors giving Raleigh its largest advantages | Strong narrative explanation of why the Triangle wins | POLOXI |
| Explains why #2/#3 lost | Explicit "Why not the others?" comparisons against Denver and Austin | Trade-offs discussed, but not systematically winner-vs-loser | POLOXI |
| Sensitivity analysis | Explicit "This ranking could change if…" naming ranking-sensitive factors | Conditional alternatives such as Huntsville if affordability matters more | POLOXI |
| Evidence sufficiency awareness | Explicitly reports LIMITED EVIDENCE | No equivalent evidence-coverage state exposed | POLOXI |
| Confidence calibration | Separates 89% match from Moderate confidence | Appropriate caveats, but no explicit confidence architecture | POLOXI |
| Match vs confidence separation | High candidate fit does not automatically mean high certainty | Not explicitly modeled | POLOXI |
| Transparency of reasoning | Query Contract → interpretations → scores → candidates → competition → uncertainty | Mostly polished narrative reasoning | POLOXI |
| Auditability | Far more of the decision process is inspectable | Hard to reconstruct exactly why Raleigh outranked every alternative | POLOXI |
| Practical consumer answer | Good, but somewhat engine-oriented | Excellent, polished, detailed relocation advice | Giant Search Engine |
| Actionability | Identifies which factors could change the result | Purchase ranges, suburbs, and concrete scouting suggestions | Giant Search Engine |
| Resistance to false certainty | Explicitly warns "directional result — limited evidence backing" | Uses caveats but still presents a confident recommendation | POLOXI |
| Follow-up reweighting support | Architecture visibly exposes factors that can be reweighted | Answers follow-ups, but no visible reusable decision structure | POLOXI |
| Enterprise explainability | Excellent | Good answer, less enterprise-auditable | POLOXI |
| Overall reasoning architecture | Multi-factor disambiguation + deterministic competition + confidence/evidence controls | Strong research-and-synthesis answer | POLOXI |
The verdict is specific, not triumphant: the Giant Search Engine produced the richer relocation research answer. POLOXI produced the more sophisticated, transparent, and auditable decision process. Candidate discovery goes to the Giant Search Engine; decision reasoning—once candidates exist—goes to POLOXI. That distinction matters: an excellent decision engine can still miss the true winner if that city never enters the candidate pool.
Where POLOXI clearly differentiated itself: "This ranking could change if…" gives the engine a representation of the decision boundary. It isn't merely saying "Raleigh is best"—it's saying "given the current interpretation and weights, Raleigh is best, and these are the variables most capable of changing that conclusion." Combined with 89% match · Moderate confidence · LIMITED EVIDENCE, POLOXI maintains three concepts that most AI answers collapse into one.
Three finalists is a narrow universe for a question spanning every U.S. city. The mathematics can be flawless and the true winner can still be missing—if Huntsville at 91% never entered the competition, Raleigh's "#1" really means "#1 among the candidates that happened to be generated." The engine itself signaled LIMITED EVIDENCE rather than presenting the result as established—exactly the behavior a reasoning architecture should exhibit—but the fix is structural:
Initial candidates: Raleigh · Austin · Denver
↓
Coverage check: national search space + only 3 candidates
↓
INSUFFICIENT → expand candidate universe
↓
Huntsville · Columbus · Madison · Charlotte · Atlanta · Minneapolis · Pittsburgh · Salt Lake City …
↓
Evidence screening → remove weak candidates → FINAL COMPETITIONThis adds a fourth headline metric alongside the existing three: Match 89% · Confidence Moderate · Evidence Limited · Candidate Coverage Moderate. A result stated that way becomes much harder to misinterpret.
One or two impressive examples aren't proof. Increasing reasoning-system complexity does not automatically improve correctness—performance depends on architecture and task characteristics. The disciplined next step is a controlled benchmark:
Ambiguity, conflicting evidence, missing evidence, misleading evidence, obvious answers, close candidates, constraint changes, follow-up reweighting, candidate omissions, stale evidence, and adversarial wording.
Base LLM vs. Base LLM + RAG vs. POLOXI—using the same underlying model wherever possible, so the architecture itself is what's being measured.
Answer correctness, constraint adherence, candidate recall, evidence precision, ranking stability, calibration, unsupported-claim rate, latency, token cost, and sensitivity correctness.
If POLOXI consistently improves those metrics, that's no longer a compelling demo—it's empirical evidence that the architecture makes the underlying LLM measurably more reliable, explainable, and useful.
Run the sequence against your own stack—then run it against POLOXI.
research@poloxi.ai →