Decision Spine
AI Analytics Reliability

Your AI analyst can write perfect SQL and still give you the wrong answer.

Decision Spine helps data teams benchmark, diagnose and fix the semantic, data-model and guardrail failures that make AI analytics silently unreliable.

Open-source research · AI Analytics Harness · Preflight

5
controlled experiments
1,488
graded runs in the repair study
48.5% → 2.4%
silent error after reliability controls
24% → 0%
wrong-metric selection after definition repair

Controlled experiments on synthetic test environments. Results demonstrate failure mechanisms, not universal model performance.

The framework

Reliable AI analytics needs three layers.

No single control makes an AI analyst trustworthy. Reliability is built in layers, cheapest and broadest first, each catching what the one below it cannot.

  • 01Broad and cheap

    Static checks

    Find ambiguity before the agent runs.

    Tool: Preflight

    • scope traps
    • concept forks
    • grain mismatch
    • definition divergence
    • duplicate or stale models
    • ambiguous metrics
  • 02Targeted and measurable

    Behavioral evals

    Test the business questions that actually matter.

    Tool: AI Analytics Harness

    • golden questions
    • repeated runs
    • metric selection
    • answerability
    • refusal
    • consistency
    • silent error
  • 03Narrow and high-stakes

    Runtime guardrails

    Stop unsafe answers before they reach users.

    • typed refusal
    • governed-tool restrictions
    • evidence requirements
    • verification
    • coverage checks

Preflight finds structural ambiguity. AI Analytics Harness measures behavior. Decision Spine applies the whole reliability stack to real data environments.

Built and led data at

Zeal GroupCoins.phNeko Health
The audit

Can your AI analyst be trusted with your numbers?

A focused assessment for data teams deploying AI analytics on top of an existing warehouse or semantic layer. Benchmark it against real business questions, find where it silently fails, and leave with a prioritized repair plan.

Approx. 10 working days · Fixed scope · From $5,000

  • Reliability benchmark

    Correct, incorrect, inconsistent, unsupported, should-refuse and silent-error behavior, measured.

  • Preflight scan

    Structural ambiguity: conflicting definitions, hidden scopes, concept forks, grain mismatches, stale or duplicate objects.

  • Failure map

    Every failure classified by source: selection, construction, coverage, ambiguity or runtime safety.

  • Prioritized repair plan

    Each issue mapped to the cheapest correct fix: model, document, declare, enforce, or remove.

  • Executive readout

    Technical report, executive summary, priority matrix and the highest-risk questions.

  • Focused retest

    Where feasible, repair one to three high-value findings, rerun the benchmark, show the before and after.

The method

Benchmark, diagnose, repair, retest.

Not a generic consulting process. A measurement loop: every claim starts from a number and ends at a re-measured one.

  1. 01

    Benchmark

    Ask realistic business questions and measure what happens.

  2. 02

    Diagnose

    Identify whether the failure comes from selection, construction, coverage, ambiguity or runtime behavior.

  3. 03

    Repair

    Fix the cheapest correct layer: modelling, semantics, documentation, declaration or enforcement.

  4. 04

    Retest

    Re-run the benchmark and measure what actually changed.

Research

AI Analytics Reliability Lab.

Controlled experiments on what makes an AI analyst reliable over business data. Change one thing at a time, hold the model and questions fixed, and measure what it buys.

All research
  • Experiment 01 · Grounding

    Question

    How much does structure buy?

    Result

    Better grounding materially improves accuracy, but capability improves faster than honesty.

    Read the experiment
  • Experiment 02 · Reliability

    Question

    Can the analyst safely say "I don't know"?

    Result

    Silent error fell from 48.5% to 2.4% across the guardrail study.

    Read the experiment
  • Experiment 03 · Protocol

    Question

    Can an answer show what its claims actually rest on?

    Result

    Citation repair improved grounded-answer coverage, while the evidence graph exposed missing reasoning structure.

    Read the experiment
  • Experiment 04 · Repair

    Question

    Which primitive failed, and where should it be fixed?

    Result

    1,488 graded runs produced a repair matrix mapping failure types to the cheapest correct layer.

    Read the experiment
  • Experiment 05 · Ambiguity

    Question

    Can we predict wrong choices before the agent runs?

    Result

    In the governed-metric subset, wrong-metric selection dropped from 24% to 0 after ambiguity repair.

    Read the experiment
Open source

Tools for measuring and preventing silent failures.

The research runs on tooling anyone can read, run and check. Both projects are open source.

  • d-n-ust/ai-analytics-harness

    AI Analytics Harness

    A reproducible lab for reliable AI analysts.

    Five controlled experiments across grounding, reliability, protocol, repair and ambiguity.

  • d-n-ust/preflight-analytics

    Preflight

    Catch ambiguous analytics before AI does.

    Static analysis for metrics, models and semantic definitions that can cause an AI analyst to choose the wrong valid object.

Want to run this against your own analytics environment?

Book a call
Dmitry Ustimov

Dmitry Ustimov

Limassol, Cyprus

Connect on LinkedIn
Who you'll work with

Built by someone who has operated the systems.

I'm Dmitry Ustimov, a Data & AI Architect with 15+ years building data platforms, analytics systems and ML infrastructure. I've led data organizations in fintech, built systems at 100TB to petabyte scale, and worked across analytics, governance, real-time services and executive decision systems.

Today I focus Decision Spine on one problem: making AI analytics reliable enough to use for real business decisions. I build the research in public through the AI Analytics Harness and Preflight, then apply those findings to real data environments.

About Dmitry
Get started

Find out where your AI analytics can silently fail.

Benchmark your AI analyst against real business questions, see the failure modes before broader rollout, and leave with a prioritized repair plan.