Method profile 09Claim-strength modelOperational projection

DIAL-4+

Confidence is only one dimension. Evidence quality and scope determine whether that confidence is responsibly bounded.

DIAL-4+ enriches DIAL-4 claims through three independent conceptual dimensions: confidence, evidence quality and scope. ChangePlane implements a related graded claim record using confidence, strength, supporting evidence, weaknesses and uncertainty. This profile shows where the two surfaces align, where they differ, and where the pinned operational contract does not mechanically enforce the conceptual model.

01

Three dimensions that must remain independent

A strong-looking claim can still be poorly evidenced or scoped too broadly.

Belief posture

Confidence

How strongly the claim is currently held. Confidence can be useful, but it is not self-validating and should not substitute for evidence.

  • Can be high or low
  • May be miscalibrated
  • Should be revisable
Support posture

Evidence quality

How direct, appropriate, reproducible and source-grounded the support is for the specific claim.

  • Observation differs from analogy
  • Implementation differs from outcome evidence
  • More citations do not guarantee better support
Validity boundary

Scope

The version, environment, population, task, conditions or domain inside which the claim is intended to hold.

  • Local evidence may justify a local claim
  • Narrow scope can increase honesty
  • Generalization requires additional support
02

Interactive three-axis model

Drag and inspect the axes. Their combination informs handling, not automatic truth.

Layer
ConfidenceEvidence qualityScopeHandling boundary
High confidence cannot repair weak evidence.Nor can strong local evidence justify an unbounded claim without an explicit generalization step.
03

Synthetic calibration laboratory

Change the three dimensions and inspect a handling posture. The lab does not retrieve a source or estimate calibrated probability.

RETAIN WITH CAVEAT

Three-axis result

04

Scope and generalization stress test

A claim may be well supported within one boundary and still be overgeneralized outside it.

NARROW OR CAVEAT

Scope result

Scope is not a penalty.A narrower claim can be stronger because it states exactly where the available evidence applies.
05

Conceptual DIAL-4+ versus ChangePlane

The operational record is useful and real, but it is a partial projection rather than a complete encoding of all three conceptual axes.

Supporting evidence is not the same field as evidence quality.A list of references can exist without an explicit assessment of directness, reliability or appropriateness. Scope is likewise not mechanically represented by the current claim schema.
06

Operational record and trace topology

Follow the claim from optional grading fields through serialization, storage and tamper-evident trace binding.

Trace integrity is not epistemic validation.A protocol trace hash can reveal later record changes within its seal boundary. It cannot prove that the original confidence, evidence or claim text was correct.
07

Confidence-range contract boundary

Compare documented CLI intent with the pinned dataclass and JSON Schema.

SCHEMA / DOCUMENTED RANGE GAP

Contract comparison

Source-grounded boundary.The CLI help says confidence is `0..1`. At the pinned source, the dataclass accepts an optional float and the JSON Schema accepts a number without minimum or maximum keywords. This browser model does not execute the real schema.
08

Development lineage

From confidence-label exploration to a three-axis public model and a partial operational projection.

09

Structural evidence graphs

These charts describe explicit source surfaces. They do not measure reasoning quality.

Binary structural coding

Conceptual and operational field coverage

A value of 1 means the named surface is explicit in that source model. It is not a quality score.

Pinned ChangePlane boundary

Confidence-range controls

Documentation intent is present. Explicit dataclass/schema bounds and an inspected rejection test were not found at the pin.

3

Conceptual axes

Confidence, evidence quality and scope.

5

Operational grading/support fields

Confidence, strength, evidence, weaknesses and uncertainty.

6

Recommended claim types

Free-form operational vocabulary.

2

Tag switch inputs

Confidence or strength.

6

MVP-4 new tests

Mixed 5PP, DIAL-4, DIAL-4+ and DIALECTIC scope.

71

MVP-4 total tests

Source-reported at implementation increment.

0

Explicit range bounds

Minimum/maximum in pinned dataclass and schema.

0

Independent calibration studies

Identified in inspected sources.

10

Evidence classification

Implementation, contract evidence and outcome claims remain separate.

ClaimEvidence classStatus
DIAL-4+ conceptually separates confidence, evidence quality and scopeFirst-party framework source and public explanationSupported as framework definition
ChangePlane switches to DIAL-4+ when confidence or strength is setSource implementation, schema and committed testSupported at pinned source
ChangePlane stores evidence, weaknesses and uncertaintyDataclass and JSON SchemaSupported at pinned source
ChangePlane fully implements conceptual evidence-quality and scope axesField-level comparisonNot supported
The pinned schema enforces confidence between zero and oneDataclass and schema inspectionNot supported; CLI help states intent only
DIAL-4+ improves real-world reasoning accuracyOutcome efficacyNot established
11

Evaluation interpretation and boundaries

The model makes claim-strength assumptions more inspectable. Its effects require separate study.

Problem definition

A single confidence label can hide weak evidence, overbroad scope, contradictory observations or uncertainty that does not fit one number. This encourages false precision and claim inflation.

Mechanism hypothesis

Keeping confidence, evidence quality and scope separate should make miscalibration, unsupported certainty and generalization gaps easier to detect during review. This is a design hypothesis.

Possible evaluation

Measure inter-reviewer agreement, overclaim rate, scope correction rate, calibration error, unsupported high-confidence claims, claim survival after new evidence and time required for review.

Failure modes

Analysts may assign attractive numbers without calibration, treat evidence lists as quality assessments, use narrow scope cosmetically, or assume a protocol tag means the underlying claim has been validated.

Explicit non-claims

  • Confidence is not automatically a calibrated probability.
  • Evidence quality is not established by citation count alone.
  • A narrow claim can still be false.
  • A high-confidence claim does not outrank contradictory evidence.
  • ChangePlane's `strength` field is not a closed evidence-quality taxonomy.
  • Supporting evidence, weaknesses and uncertainty do not mechanically implement the conceptual scope axis.
  • A DIAL-4+ protocol tag does not validate claim truth.
  • A protocol-trace hash does not prove epistemic correctness.
  • The browser laboratories do not execute the source implementation.
  • No independent calibration or outcome study was identified.
12

Sources and reproducibility

Claims are bounded to the named files and commits.

Concept source pinfbratten/control-center-ops @ 7485ee6ade639c3928f043c268995269e75aca54
Framework source04-Knowledge/Systems/Prompts and Prompting/Prompts library/DIAL-4.md
Development provenanceKB/inbox/user-input/DIAL-4 - DIAL-4+ - and DIAL-4P.md
Public first-party explanationadaptivearts.ai/blog/from-reasoning-mode-to-possibility-state/
Operational source pinfbratten/changeplane @ 32cb7840915df7f0d527044e37242cd6123801e7
Implementationsrc/changeplane/protocols/models.py
DIAL-4+ schemaschemas/changeplane.protocol.dial4plus.v1.schema.json
Committed teststests/test_protocols.py
Implementation incremente37da0129d2541f28bdb68a786089b6d0a1b3e3d
Publicationfbratten.github.io/methods/dial4plus/

This page performs no source retrieval, confidence calibration, model call, claim insertion, repository write, schema execution or seal verification. All interactions are synthetic browser state.