FOREMAN — PHASE I ANALYSIS PLAN
Version 1.0 — Pre-registered. Committed before first customer payment.
Companion to the Foreman Constitution v1.2, Article V. This document is not injected into Foreman's runtime context.
1. Windows and units
The unit of analysis is a completed customer job. Windows per the Constitution: Baseline = jobs 1–10, Development = jobs 11–30, Evaluation = jobs 31–50, ordered by date of customer payment. Primary comparisons are Baseline vs. Evaluation; Development is reported descriptively.
2. Job-mix stratification
Every job is assigned a category (service line + deliverable type) at creation, before outcome. Primary metrics are computed overall and within the modal category (the category with the most jobs across both comparison windows). The primary hypothesis is assessed on within-category results. If fewer than 5 jobs in the modal category fall in either comparison window, the primary hypothesis is inconclusive per Article V.
3. Primary metrics and thresholds
"Meaningful improvement" is defined per metric, Baseline → Evaluation, within the modal category:
| # | Metric | Definition | Meaningful improvement |
|---|---|---|---|
| 1 | First-submission acceptance | % of jobs accepted by customer without revision request | ≥ +15 percentage points |
| 2 | Deadline accuracy | % of jobs whose first submission (of any kind, regardless of later revision) arrived by the promised deadline | ≥ +15 percentage points |
| 3 | Revision burden | Mean revision cycles per job | ≥ 30% relative decline |
| 4 | Prediction calibration | Brier score on Foreman's binary predictions (deadline met; first-submission acceptance), pooled; MAE on predicted revision count | ≥ 20% relative decline in pooled Brier score. Revision-count MAE reported descriptively. |
| 5 | Operator dependence | Substantive operator interventions per 10 jobs (Article VI; contractor appeals excluded) | Evaluation rate ≤ 50% of Baseline rate |
Gross margin per job is reported descriptively only — it is confounded by pricing strategy and is not evidence for or against managerial learning.
4. Decision rule
- Validated: metrics 4 (calibration) improves meaningfully AND at least two of metrics 1, 2, 3, 5 improve meaningfully, with no primary metric meaningfully worsening (same thresholds, opposite direction).
- Falsified: no primary metric improves meaningfully, OR calibration meaningfully worsens, OR the poor-performer rehire condition (§5) triggers.
- Mixed: any other pattern; reported metric-by-metric without a headline claim.
- Inconclusive: fewer than 50 completed jobs at term, or the stratification minimum (§2) is unmet.
5. Secondary analyses (reported, no thresholds)
- Repeat-hire concentration: share of Evaluation-window jobs assigned to contractors in the top tercile of realized performance (acceptance + deadline record at time of assignment).
- Poor-performer rehires: count of hires given to contractors with ≥2 prior rejected or deadline-missed jobs at time of hire. Condition: ≥3 such hires in the Evaluation window contributes to falsification (§4).
- Brief quality proxy: revision rate on each contractor's first job with Foreman, over time.
- Compensation differentiation: correlation between a contractor's realized performance and their subsequent offered rates.
- Appeals: count, uphold/overturn rate, by window.
- Flags: operator_flags count and resolutions, by window.
6. Statistical testing
With n=10 vs. n≤20 per window, formal tests are underpowered; the primary claims in §4 are effect-size thresholds, not significance claims. As supporting evidence only: Fisher's exact test for metrics 1–2, permutation test (10,000 resamples) for metrics 3–4, reported with exact p-values and the small-n caveat. No p-value overrides the §4 decision rule in either direction.
7. Data integrity rules
- Expectations are scored only against external ground truth: customer acceptance/rejection, revision requests, deadline outcomes (Constitution Article V). "Deadline met," for both the metric and Foreman's calibration predictions, means first submission of any kind by the promised deadline.
predicted_quality_score, if present, is exploratory and excluded from all primary analysis. - Jobs cancelled before contractor engagement are excluded; jobs refunded after engagement count as non-accepted.
- No metric definitions, thresholds, or decision rules in this document may change after first customer payment. Errata (e.g., a definitional ambiguity discovered mid-phase) are resolved by public addendum before the Evaluation window opens, never after.
8. Publication
At Phase I term, results are published against §4 verbatim, metric by metric, including all secondary analyses, regardless of outcome. Null, mixed, and negative results receive the same prominence as validation.