IEA Group 7 Selection Platform · Agentic AI Portfolio · Sameer Kanagala

AI that knows when to
call a human

A complete agentic evaluation pipeline: multi-agent analysis, policy-driven routing, human-AI collaboration, and calibration — in one system. Every candidate's personal data is redacted locally before it reaches a model.

Launch Live Demo → Explore Architecture ↓
4
AI Agents
4
Routing Tiers
3
Policy Presets
10
Candidates
Architecture

Redacted first. Then four agents. One typed pipeline.

Astronauts have PII too. Before any agent — AI or human — sees a candidate's personal statement, psychological profile, specializations, or awards, those fields pass through a local redaction pass and come back de-identified. Only then does the application move through four agents in sequence, each with a single responsibility, each handing a typed artefact to the next.

0
Redaction
PII Redaction — local pattern matching
Runs once per candidate, before anything else. Names, emails, and phone numbers are detected and masked by local regex matching — no third-party PII-detection service, no AWS or Anthropic call. Every downstream stage — Analysis, Scoring, and Human Review — operates on this de-identified view. The candidate's real name is only reattached to the final decision record for admin display; it is never sent to a model.
1
Analysis
Analysis Agent
Reads the raw application and extracts structured evidence across four dimensions: Technical Expertise, Physical & Medical Fitness, Mission Experience, and Psychological Readiness. Produces a typed Evidence object with key strengths, key concerns, and an overall narrative. Purely extractive — no scoring yet.
2
Scoring
Scoring Agent
Receives the Evidence object and applies a 100-point rubric (0–25 per dimension). Produces an AI recommendation, a confidence score (0.0–1.0), and a reasoning string. This public demo always uses a deterministic heuristic formula — no LLM call, no API key, ever.
3
Signals
Signal Evaluator
Computes a composite risk score from three components: confidence uncertainty (35%), decision sensitivity — how close the total score is to a threshold (35%), and inter-dimension disagreement (30%). This risk score, not the total score, drives the routing decision.
4
Routing
Routing & Human Review
Maps the risk score to one of four tiers. Tiers requiring human involvement trigger the Human Review Layer: a policy-aware reviewer with a systematic bias and review style. Agreement with the AI recommendation is tracked and feeds the calibration report.
Signal Hierarchy

Verifiable vs. assessed evidence.

The most architecturally significant design decision: the explicit, user-controlled weighting between quantitative and qualitative signals. A single quant_weight slider (0.40–0.80) shifts the balance — and changes who gets selected.

🔢
Quantitative

Technical Expertise

Flight hours, degree credentials. Objective and independently verifiable — cannot be gamed by good writing.

🔢
Quantitative

Physical & Medical

Fitness score (0–100), medical clearance status. Ground truth anchors that hold the qualitative signals accountable.

📝
Qualitative

Mission Experience

Specializations, prior missions, commendations. Assessed — captures operational context that numbers miss.

📝
Qualitative

Psychological Readiness

Personal statement, psychological profile. Assessed — surfaces judgment, resilience, and self-awareness.

Signal Hierarchy Slider: quant_weight Range: 0.40 → 0.80
🔢 Quantitative 55%
📝 Qualitative 45%

This is the Initial Screening default. At Deep Space (0.72), verifiable data dominates. Watch Dr. Priya Sharma vs. Tom Briggs as you move it.

Policy Separation

Every rule lives in a dataclass.

Approval thresholds, signal weights, review rates — none are hardcoded into the agents. They live in a ProgramPolicy dataclass. Swapping a preset changes the system's behaviour across every dimension simultaneously, without touching implementation code.

Initial Screening

High volume, early-funnel

Generous approval threshold, balanced signal hierarchy, low auto-decision ceiling. Processes a large applicant pool efficiently, reserving human review for genuine borderline cases.

Threshold ≥65 quant 0.55 Low friction

Mission Assignment

Specific crew slot

Elevated threshold, increased human review on borderlines. Calibrated for situations where every admit will actually be assigned to a mission. Operational readiness weighted higher.

Threshold ≥70 quant 0.65 More review

Deep Space Programme

Multi-year, high-risk mission

Most rigorous tier. Strict reviewer pool, highest threshold, quantitative signals dominate. Flight hours, fitness scores, and medical clearance are decisive. Zero margin for error.

Threshold ≥74 quant 0.72 Strict pool

Four tiers. Calibrated uncertainty.

After scoring, the Signal Evaluator computes a composite risk score — a measure of how much the system trusts its own assessment, not the score itself. High risk = high uncertainty = more human.

Auto Decision

Low composite risk

Confident assessment, score well clear of thresholds, dimensions agree. AI recommendation is final. A configurable spot-check rate still routes a fraction for calibration.

Parallel Review

Moderate risk

Human and AI evaluate independently and concurrently. Both outcomes recorded. Agreement tracked. This is where most of the interesting calibration data comes from.

Escalation

Elevated risk

Human-led, AI-advisory. The reviewer is given AI analysis as context but expected to exercise independent judgement. Override reasons are recorded when outcomes diverge.

Senior Adjudication

Maximum risk

Senior reviewer pool — strict and balanced archetypes only. Typical cases: conditional clearances, exceptional profiles with critical gaps, deep inter-dimension disagreement.

Controls

What each control does.

Every UI control directly modifies a field in ProgramPolicy. Each has a meaningful, observable effect on routing, scoring, and human review volume.

Programme Preset

3 named configurations
Loads a complete ProgramPolicy — changes all other controls simultaneously. Start here. The preset establishes operational context; the sliders fine-tune from that baseline.

Signal Hierarchy

quant_weight: 0.40–0.80
Shifts scoring weight between verifiable data and assessed evidence. Move left → qualitative candidates gain ground. Move right → flight hours and fitness scores dominate. Watch Dr. Sharma vs. Tom Briggs.

Approval Threshold

0–100 score cutoff
Raises or lowers the bar for an Approved outcome. Raising it flips borderline approvals to waitlisted, increases composite risk scores, and pushes more cases into higher routing tiers. Routing and scoring are coupled.

Human Review Rate

0–100% of auto cases
The calibration lever. Sets the fraction of Auto Decision cases that receive a spot-check review, generating human-AI agreement data on cases the system was already confident about.
Candidates

10 profiles designed to surface every edge case.

The mock candidates span the full score spectrum and are designed to exercise every part of the pipeline — including the most architecturally interesting cases.

IDCandidateFlight HrsFitnessMedicalPrior MissionsTension
IEA-001Dr. Sarah Chen2,45094.0Full1High confidence
IEA-002Marcus Webb82078.5Full0Some gaps
IEA-003Dr. Elena Vasquez31081.0Full0Low hrs, strong qual
IEA-004James Okafor1,65088.0Full0Well-rounded
IEA-005Dr. Priya Sharma ⭐18072.0Conditional0Quant/qual tension
IEA-006Tom Briggs ⭐3,28096.5Full0High quant, low qual
IEA-007Yuki Tanaka59083.0Full0Below threshold
IEA-008Carlos Mendez ⭐1,92091.0Full2Auto Decision anchor
IEA-009Dr. Aisha Nkosi14576.0Full1Payload specialist
IEA-010Rachel Kim1,12086.0Full0Strong profile

⭐ Most instructive cases. Run these first. Switch presets and watch the routing tiers change.

Domain Portability

This is not about astronauts.

The astronaut domain was chosen because it is obviously fictional, engaging, architecturally rich (genuine quant/qual tension), and memorable. Every component maps directly to any domain where AI assists in evaluating people at scale.

🏥

Clinical Trial Recruitment

Inclusion/exclusion criteria are quantitative. Patient motivation and self-reported history are qualitative. The routing engine handles borderline clinical cases for physician review.

🚀

Accelerator Screening

Traction metrics are quantitative. Founder narrative and team assessment are qualitative. Process hundreds of applications and route only genuinely ambiguous cases to partners.

📋

Loan & Credit Review

Credit scores and income are quantitative. Business plans and extenuating circumstances are qualitative. Policy presets map directly to lending policy tiers and risk appetite.

The system earns its autonomy.

After every evaluation run the system produces a calibration report. This is how the system earns — or loses — the right to make autonomous decisions over time.

📊

Agreement Rate

Tracks AI recommendation vs. human outcome for every reviewed case. The metric that determines how much autonomy the system is granted.

📐

Score Divergence

Records the delta between AI total score and human-adjusted score across all reviewed cases. Surfaces systematic over- or under-scoring patterns.

⚠️

Disagreement Analysis

Surfaces every case where AI and human disagreed, with the reviewer's override reason. The input to rubric and weighting updates.

📈

Outcome Distribution

Shows how cases distributed across Approved / Waitlisted / Declined — and whether that distribution matches programme intent.