Confidence
How strongly the claim is currently held. Confidence can be useful, but it is not self-validating and should not substitute for evidence.
- Can be high or low
- May be miscalibrated
- Should be revisable
Confidence is only one dimension. Evidence quality and scope determine whether that confidence is responsibly bounded.
DIAL-4+ enriches DIAL-4 claims through three independent conceptual dimensions: confidence, evidence quality and scope. ChangePlane implements a related graded claim record using confidence, strength, supporting evidence, weaknesses and uncertainty. This profile shows where the two surfaces align, where they differ, and where the pinned operational contract does not mechanically enforce the conceptual model.
A strong-looking claim can still be poorly evidenced or scoped too broadly.
How strongly the claim is currently held. Confidence can be useful, but it is not self-validating and should not substitute for evidence.
How direct, appropriate, reproducible and source-grounded the support is for the specific claim.
The version, environment, population, task, conditions or domain inside which the claim is intended to hold.
Drag and inspect the axes. Their combination informs handling, not automatic truth.
Change the three dimensions and inspect a handling posture. The lab does not retrieve a source or estimate calibrated probability.
A claim may be well supported within one boundary and still be overgeneralized outside it.
The operational record is useful and real, but it is a partial projection rather than a complete encoding of all three conceptual axes.
Follow the claim from optional grading fields through serialization, storage and tamper-evident trace binding.
Compare documented CLI intent with the pinned dataclass and JSON Schema.
From confidence-label exploration to a three-axis public model and a partial operational projection.
These charts describe explicit source surfaces. They do not measure reasoning quality.
A value of 1 means the named surface is explicit in that source model. It is not a quality score.
Documentation intent is present. Explicit dataclass/schema bounds and an inspected rejection test were not found at the pin.
Confidence, evidence quality and scope.
Confidence, strength, evidence, weaknesses and uncertainty.
Free-form operational vocabulary.
Confidence or strength.
Mixed 5PP, DIAL-4, DIAL-4+ and DIALECTIC scope.
Source-reported at implementation increment.
Minimum/maximum in pinned dataclass and schema.
Identified in inspected sources.
Implementation, contract evidence and outcome claims remain separate.
| Claim | Evidence class | Status |
|---|---|---|
| DIAL-4+ conceptually separates confidence, evidence quality and scope | First-party framework source and public explanation | Supported as framework definition |
| ChangePlane switches to DIAL-4+ when confidence or strength is set | Source implementation, schema and committed test | Supported at pinned source |
| ChangePlane stores evidence, weaknesses and uncertainty | Dataclass and JSON Schema | Supported at pinned source |
| ChangePlane fully implements conceptual evidence-quality and scope axes | Field-level comparison | Not supported |
| The pinned schema enforces confidence between zero and one | Dataclass and schema inspection | Not supported; CLI help states intent only |
| DIAL-4+ improves real-world reasoning accuracy | Outcome efficacy | Not established |
The model makes claim-strength assumptions more inspectable. Its effects require separate study.
A single confidence label can hide weak evidence, overbroad scope, contradictory observations or uncertainty that does not fit one number. This encourages false precision and claim inflation.
Keeping confidence, evidence quality and scope separate should make miscalibration, unsupported certainty and generalization gaps easier to detect during review. This is a design hypothesis.
Measure inter-reviewer agreement, overclaim rate, scope correction rate, calibration error, unsupported high-confidence claims, claim survival after new evidence and time required for review.
Analysts may assign attractive numbers without calibration, treat evidence lists as quality assessments, use narrow scope cosmetically, or assume a protocol tag means the underlying claim has been validated.
Claims are bounded to the named files and commits.
fbratten/control-center-ops @ 7485ee6ade639c3928f043c268995269e75aca5404-Knowledge/Systems/Prompts and Prompting/Prompts library/DIAL-4.mdKB/inbox/user-input/DIAL-4 - DIAL-4+ - and DIAL-4P.mdadaptivearts.ai/blog/from-reasoning-mode-to-possibility-state/fbratten/changeplane @ 32cb7840915df7f0d527044e37242cd6123801e7src/changeplane/protocols/models.pyschemas/changeplane.protocol.dial4plus.v1.schema.jsontests/test_protocols.pye37da0129d2541f28bdb68a786089b6d0a1b3e3dfbratten.github.io/methods/dial4plus/This page performs no source retrieval, confidence calibration, model call, claim insertion, repository write, schema execution or seal verification. All interactions are synthetic browser state.