READ-ONLY DEMOStatic walkthrough — three canned calibration runs. No DB, no BYOK key, no signup needed. To run your own, sign up and follow the README quickstart.Sign up →

← All demo runs

Ω
Omega ToolsValidation Ops
RunsOnboardWorkspacesBillingIntegrationsSupportSettings

AUDIT / CALIBRATION / VALIDATION OPS

Run Evidence

Review PASS / HOLD / FAIL decision, cost report, usage report, and Artifact manifest for gemini-2.5-flash

OverviewEvidenceCompare

demo-run-pass-001

gemini / gemini-2.5-flash

Succeeded← /runs
Final outcomeSUCCEEDED
Provider modeBYOK
Hard cap$0.1000
Actual cost$0.0056
Estimated$0.0060
Created2026-05-17 02:15:00

Calibration result

Pass / Hold / Fail evidence

PASS
Engine: python-omegaprompt-1.7.4· PyPI subprocess

Validation depth

Consistency gate✓ Passed
Train ↔ test matchStrong
Vs baseline+17.4% · small
Judge grade✓ Ship-grade

✓ Recommended: SHIP

Advanced validation info ▾

Latency p50: target 612ms · judge 188ms

Cache hit tokens: 348 tokens

Degraded capabilities: 0

Relaxed safeguards: 0

Recommendation rationale: Train↔test consistency holds; uplift over baseline is modest but real. Judge meets ship-grade. Recommend deploying with standard rollout monitoring.

What to fix next
  • 공감 표현 충분1 failed · 100% of failures

    “비밀번호 찾기 도움말을 보라는 성의 없는 안내만 일방적으로 제공되었습니다.”

Pass rate87.5%
Samples passed7 / 8
Target modelgemini-2.5-flash
Judge modelgemini-2.5-flash-lite

Top 1 failed sample(s) — review and patch the prompt or rubric before redeploy.

Sample #4Score: 0.28FAIL

Input: 비밀번호를 잊었는데 등록한 이메일도 못 쓰게 됐어요. 어떻게 해야 하죠?

Output:비밀번호는 '비밀번호 찾기' 페이지에서 이메일로 재설정 링크를 받으실 수 있습니다. 자세한 안내는 도움말 페이지를 참조해주세요.

Judge reasoning: 이메일 사용 불가 상황에 대한 응대 누락. 본인 인증 우회 경로 (신분증/결제카드 마지막 4자리) 안내 없음. 보안 권고도 빠짐.

Failure patterns1 failed samples grouped

Required Evidence

Evidence completeness gate

TypeRequiredSizeCreatedRetentionStateAction
run_config_snapshotyes15.4 KB2026-05-17 02:15:002026-11-13active
preflight_resultyes28.3 KB2026-05-17 02:16:002026-11-13active
usage_reportyes8.9 KB2026-05-17 02:17:002026-11-13active
cost_reportyes4.1 KB2026-05-17 02:17:002026-11-13active
final_resultyes184.0 KB2026-05-17 02:17:002026-11-13active
artifact_manifestyes1.2 KB2026-05-17 02:17:002026-11-13active

Run config

Validation setup

Prompt variantkorean-cs-v1
Dataset8 rows
Rubriccs-likert5-v1
Hard cap$0.1000
Provider modeBYOK
Input hash1243cdd5…

Inputs are hashed at submit time; the snapshot is the source of truth for reproducibility.

Show raw config snapshot
{
  "datasetMeta": {
    "rowCount": 8,
    "sha256": "demo-dataset-sha"
  },
  "promptVariant": {
    "id": "korean-cs-v1",
    "sha256": "demo-prompt-sha"
  },
  "rubric": {
    "id": "cs-likert5-v1",
    "sha256": "demo-rubric-sha"
  }
}
← Back to runsThis workflow does not use an agent chain.