Experiment 01 · Grounding
Question
How much does structure buy?
Result
Better grounding materially improves accuracy, but capability improves faster than honesty.
Decision Spine helps data teams benchmark, diagnose and fix the semantic, data-model and guardrail failures that make AI analytics silently unreliable.
Open-source research · AI Analytics Harness · Preflight
Controlled experiments on synthetic test environments. Results demonstrate failure mechanisms, not universal model performance.
No single control makes an AI analyst trustworthy. Reliability is built in layers, cheapest and broadest first, each catching what the one below it cannot.
Find ambiguity before the agent runs.
Tool: Preflight
Test the business questions that actually matter.
Tool: AI Analytics Harness
Stop unsafe answers before they reach users.
Preflight finds structural ambiguity. AI Analytics Harness measures behavior. Decision Spine applies the whole reliability stack to real data environments.
Built and led data at
A focused assessment for data teams deploying AI analytics on top of an existing warehouse or semantic layer. Benchmark it against real business questions, find where it silently fails, and leave with a prioritized repair plan.
Approx. 10 working days · Fixed scope · From $5,000
Correct, incorrect, inconsistent, unsupported, should-refuse and silent-error behavior, measured.
Structural ambiguity: conflicting definitions, hidden scopes, concept forks, grain mismatches, stale or duplicate objects.
Every failure classified by source: selection, construction, coverage, ambiguity or runtime safety.
Each issue mapped to the cheapest correct fix: model, document, declare, enforce, or remove.
Technical report, executive summary, priority matrix and the highest-risk questions.
Where feasible, repair one to three high-value findings, rerun the benchmark, show the before and after.
Not a generic consulting process. A measurement loop: every claim starts from a number and ends at a re-measured one.
Ask realistic business questions and measure what happens.
Identify whether the failure comes from selection, construction, coverage, ambiguity or runtime behavior.
Fix the cheapest correct layer: modelling, semantics, documentation, declaration or enforcement.
Re-run the benchmark and measure what actually changed.
Controlled experiments on what makes an AI analyst reliable over business data. Change one thing at a time, hold the model and questions fixed, and measure what it buys.
Experiment 01 · Grounding
Question
Result
Better grounding materially improves accuracy, but capability improves faster than honesty.
Experiment 02 · Reliability
Question
Result
Silent error fell from 48.5% to 2.4% across the guardrail study.
Experiment 03 · Protocol
Question
Result
Citation repair improved grounded-answer coverage, while the evidence graph exposed missing reasoning structure.
Experiment 04 · Repair
Question
Result
1,488 graded runs produced a repair matrix mapping failure types to the cheapest correct layer.
Experiment 05 · Ambiguity
Question
Result
In the governed-metric subset, wrong-metric selection dropped from 24% to 0 after ambiguity repair.
The research runs on tooling anyone can read, run and check. Both projects are open source.
AI Analytics Harness
Five controlled experiments across grounding, reliability, protocol, repair and ambiguity.
Preflight
Static analysis for metrics, models and semantic definitions that can cause an AI analyst to choose the wrong valid object.
I'm Dmitry Ustimov, a Data & AI Architect with 15+ years building data platforms, analytics systems and ML infrastructure. I've led data organizations in fintech, built systems at 100TB to petabyte scale, and worked across analytics, governance, real-time services and executive decision systems.
Today I focus Decision Spine on one problem: making AI analytics reliable enough to use for real business decisions. I build the research in public through the AI Analytics Harness and Preflight, then apply those findings to real data environments.
Benchmark your AI analyst against real business questions, see the failure modes before broader rollout, and leave with a prioritized repair plan.