Content Factory • Methodology v0.1
A task’s “AI vulnerability” is not one number. We need to know what AI can execute, what still requires a human owner, how much of the workflow is exposed to automation, and whether the task is documented well enough to test.
The Content Factory AI Capability Index turns the Task Library into a living evaluation system. Its purpose is to route work responsibly, improve skill files, train outcome owners, and upgrade evidence as real attempts succeed or fail. It does not predict the probability that a person or job will be replaced.
Four measures, not a fake IQ
The index evaluates a task under a named stack: model, reasoning mode, tools, memory, permissions, organizational context, acceptance test, and reviewer. Change the stack and the result can change.
| Measure | Question answered | What it must not be mistaken for |
|---|---|---|
| AI Execution Capability | Can the defined AI stack produce an acceptable task artifact? | General intelligence or production reliability |
| Human Accountability Demand | How much judgment, relationship ownership, approval, or consequence-bearing must remain human? | The percentage of labor performed manually |
| Automation Exposure | How much of the defined workflow could plausibly be automated after integration and controls? | Job-loss probability or realized adoption |
| Documentation Readiness | Are the outcome, inputs, skill, access, acceptance criteria, and failure handling explicit enough to run a fair test? | Proof that the task works in production |
What the current scores mean
Version 0.1 is an E0 triage model. It uses task category and language signals—such as physical-world work, relationship dependence, deterministic steps, external actions, spending, high-stakes consequences, long horizons, and privileged access—to generate a low-confidence starting hypothesis. It is deliberately inexpensive so all registered tasks can be ranked for testing.
The intended measured version should compute a Task Capability Index from five normalized trial results:
TCI = 100 × (Quality × Reliability × Autonomous Share × End-to-End Scope × Generalization)1/5
This formula describes the intended measured methodology; it is not the heuristic currently generating E0 dashboard scores. A separate deployment coefficient should then account for integration, permission boundaries, verification and recovery, accountable ownership, and economics. Multiplying measured capability by that coefficient can estimate automation viability. It still does not estimate displacement.
Non-negotiable guardrails
- Every score is conditional on a dated task version and tested AI stack.
- Capability, exposure, accountability, documentation, and evidence remain separate fields.
- A task does not receive a production-capable label from a polished demonstration alone.
- Higher-risk permissions require stronger evidence, tighter thresholds, and a named human owner.
- Model upgrades can improve a result; tool, prompt, access, or environment changes can also break it.
- Scores must be re-tested after material workflow changes and marked stale when evidence expires.
The E0–E4 evidence ladder
| Level | Evidence | Allowed claim |
|---|---|---|
| E0 | Heuristic or expert hypothesis; no controlled attempt | Worth testing |
| E1 | A working demonstration on a known example | Can work in this case |
| E2 | Repeated or held-out cases with recorded acceptance results | Generalizes within the tested range |
| E3 | Human-only, AI-only, and human-plus-AI comparison against a credible baseline | Measured relative performance |
| E4 | Sustained production outcomes, incidents, cost, recovery, and downstream effects | Production evidence under stated conditions |
Three Content Factory tasks show why separation matters
| Task | Expected score shape | Human responsibility |
|---|---|---|
| Record an authentic one-minute story | Low execution capability for the lived experience itself; AI can coach, outline, transcribe, and diagnose. | The person supplies truth, consent, voice, and relationship credibility. |
| Draft platform posts from an approved transcript | High execution capability and exposure when voice rules and acceptance tests are supplied. | An owner verifies accuracy, claims, context, and fit before publication. |
| Launch or change a paid campaign | Potentially high execution capability but constrained automation viability because money, access, and external consequences are involved. | A named budget owner sets limits, approves launch, monitors impact, and can reverse the action. |
The operating loop is Plumbing → Produce → Process → Post → Promote → Perform. “Reason” is the cross-cutting planning and judgment layer, not a fractional production stage. The index should accelerate the loop without severing it from source truth or business outcomes.
The task-page contract
A task cannot be evaluated fairly if the page is only a title and prose. Every task page should resolve to a stable, versioned contract containing:
- Stable task ID, version, owner, and last-tested date.
- Outcome, scope, dependencies, and downstream business result.
- Required inputs, context, examples, and expected artifact.
- Skill file, tools, access, permissions, and environment.
- Acceptance criteria, critical failures, and definition of done.
- Human baseline, expected time, and task-share assumptions.
- Error detection, escalation, rollback, and recovery procedure.
- Tested model/tool stack, evidence level, confidence, and result history.
Grade students on verified outcomes—not on hiding AI use
Colleges can use Task Library work as a supervised practicum. Students should disclose the stack and decision trail, then be evaluated on the result they own. A proposed pilot rubric is: verified outcome 40%, quality and safety 20%, judgment and accountability 15%, efficiency and economics 10%, and transfer-quality documentation 15%.
Builder autonomy ladder
- L0 — Observe: study attempts, failures, and acceptance decisions.
- L1 — Sandbox: produce artifacts with no live-system effects.
- L2 — Assisted: use read-only access or submit every output for review.
- L3 — Review-gated live: perform reversible actions after explicit approval.
- L4 — Bounded owner: act within documented thresholds, budgets, and escalation rules.
- L5 — Architect and mentor: design the workflow, evaluate others, and improve its controls.
Autonomy cannot exceed the lower of the evidence ceiling and the risk approval. E0 and E1 do not justify unsupervised live action. Credentials, DNS, billing, sensitive data, high-stakes claims, and irreversible actions always require a named accountable owner and explicit controls.
The reward loop should favor pre-release error detection instead of hiding mistakes. Track Verified Outcome Achievement Rate, supervision minutes per accepted outcome, and the sustained paid outcome-owner rate. Guardrails include escaped incidents, privacy failures, inequitable access, and unpaid production work.
The candid 239 / 248 / 253 reconciliation
All three numbers appeared in BlitzMetrics materials, and all three described different universes in the August 2, 2026 audit snapshot:
- 239 = local hub task-skill files maintained in the repository snapshot.
- 248 = runnable skills in the older v3.3 WordPress bundle: 239 task files plus nine top-level operational skills. That archive also contained three documentation Markdown files.
- 253 = registered records in the public dashboard snapshot, combining local hub tasks with spoke and tracker-backed entries.
The correction is not to choose the most impressive number. It is to label the universe, derive totals from the build, date the snapshot, and make the live registry the source of truth. The same audit found redirects, eight effectively empty public destinations, and many task records without a definitive article or resolvable skill link. Those are visible backlog items—not evidence of completeness.
What outside research supports—and what it does not
- Anthropic’s March 2025 Economic Index report mapped automation and augmentation behavior at task and occupation level. It did not publish a sector-by-sector probability of worker replacement.
- Anthropic’s March 2026 labor-market study introduced “observed exposure,” found no systematic unemployment increase in highly exposed occupations, and reported suggestive evidence of slower hiring among younger workers in exposed occupations.
- OpenAI’s GDPval evaluates economically valuable real-world tasks across 44 occupations and reports that reasoning effort, task context, and scaffolding materially affect performance.
- METR’s task-completion time-horizon work shows why success on a short step does not prove reliable completion of a long workflow.
- The ILO’s 2025 exposure study evaluates tasks within occupations and concludes that transformation is more likely than complete redundancy for most jobs because human input remains necessary.
- NIST’s AI Risk Management Framework supports continuous governance, mapping, measurement, management, pre-deployment testing, monitoring, recovery, and documentation.
Together, these sources support task-level, context-dependent evaluation. They do not justify converting an E0 score into a forecast of layoffs, wages, or business adoption.
How this strengthens the ten-year mission
Relationships, trust, lived expertise, distribution, and accountable business ownership are durable assets. AI can multiply them when they are connected to explicit tasks, context, access, and verification. The sharper mission is to equip people to own verified outcomes using AI labor, while measuring whether those outcomes persist and create paid opportunity.
“One million outcome owners” can be a north star. Until the denominator, time window, paid status, attribution, and verification method are defined, it must remain an aspiration rather than a reported result.
Build receipts and open limits
| Item | Result |
|---|---|
| Inventory audited | Task Library dashboard and repository, public task destinations, Content Factory materials, Spotlight Network, and the BlitzMetrics meta-article standard |
| Source changes submitted | Derived counts, stable task deep links, E0 scorecards, source-type labels, guardrail copy, and removal of hard-coded completeness claims |
| Verification completed | Local data build, Python compilation, JavaScript syntax check, whitespace check, browser rendering, modal scorecard, and direct task-link behavior |
| Known boundary | The local build produces 240 records without the private tracker feed; the deployed snapshot shows 253. Upstream PR #1 has no code conflicts and is awaiting maintainer approval of its forked workflow, review, and merge. |
| Not yet proven | No task has been upgraded merely because the E0 heuristic exists. Controlled E1–E4 evidence still has to be collected. |
| Cost receipt | Token and billing totals were not exposed as auditable telemetry in this run, so no cost number is reported or implied. |
Snapshot and methodology updated . Counts and capability labels should be regenerated, not copied forward as permanent facts.
The deliverable
Inspect tasks, scorecards, skill files, gaps, and definitive articles
Treat every E0 score as a testable hypothesis. Improve the task contract, run the work, record the error, and upgrade the evidence.
Continuation receipt · August 2, 2026 (PDT)
The final architecture separates task documentation, task scoring, and trusted evidence
Answer first: the proposed source now has three deliberately separate layers. An exact taskPage teaches a bounded task; a definitiveArticle explains the broader concept or hub; and an evidence record tests a frozen task and stack. A URL may serve both documentation roles, but neither kind of page is capability evidence. Scores belong to individual tasks, while categories report only task-score medians, distributions, modes, and evidence coverage. The current local catalog has zero validated trial records, so all 240 task records honestly remain E0. The proposed files are publicly reviewable in PR #1, but they are not merged into the base repository, active there, or deployed while the governance gate remains unresolved; the previously described WordPress repairs are the separate live-content work.
Reader path: begin with the Content Factory, browse the legacy Task Library, inspect the current Task Library Dashboard, and use the Meta-Article Prompt to see how implementation receipts should be documented. These contextual links connect the older explanations to the new scoring and evidence layer for both readers and crawlers.
Current architecture and audit baseline
- Documentation fields are not interchangeable.
taskPageidentifies the exact SOP or checklist for the bounded task.definitiveArticleidentifies the broader explanatory article or hub. The fields may point to the same URL when one page genuinely performs both jobs, but their roles, coverage counts, and audit criteria remain separate. Neither field establishes an E1–E4 result. - Exact task-page coverage: the generated task-page audit maps 39 of 240 task records to 18 unique task pages, leaves 201 records unmapped, records one legacy-alias record, and reports no mapping errors. Its current public-health pass found all 18 unique mapped task pages clean, 18/18 returning 2xx with no findings. A separate inventory of the legacy Task Library found 63 task-like final URLs: 12 map to the current catalog and 51 do not. The machine pass reported 61 clean and two Local Service Spotlight URLs with four blocking 403/non-HTML findings, but authenticated Chrome verification found both Client Onboarding and Descript Login substantive. Those four findings are recorded machine/WAF limitations, not broken-article findings. The source-index counts are a different universe from the 39-record registry mapping and are not combined.
- Definitive-article coverage: the current generated registry maps 199 of 240 task records to 17 unique definitive-article URLs and leaves 41 explicit article gaps. All 240 local records still resolve to skill Markdown. The current public-article audit checked all 17 unique URLs: 17/17 returned 2xx and were clean, with no blocking findings or warnings.
- MarketScale documentation gap resolved:
process-videos-via-marketscalenow maps both documentation fields to How to Process Videos via MarketScale. One compatibility alias points to the same canonical task and URL, so one unique page resolves two task records. This closes the documentation mapping gap, not the capability gap: both records remain needs-work and E0 until a current privacy-safe run and reviewer receipt qualify. - Task-level scoring only: each task owns
aiExecution,humanAccountability,automationExposure, andreadiness. A category reports medians, distributions in the fixed 0–19, 20–39, 40–59, 60–79, and 80–100 bands, mode counts, and evidence coverage split between E0/E1–E4 and by evidence level. It does not receive a synthetic replacement score. Automation exposure is a task-level E0 heuristic, not a probability that a job will be replaced and not an employment forecast. - Directional link graph: the generated graph audit covers 39 mapped task records and 27 unique fetched pages, with zero fetch failures. It found 9 task-page→library links, 0 task-page→methodology links, 34 task-page→definitive-article links, 25 definitive-article→task-page links, and 5 complete context loops. These are directional record-level coverage counts; direct methodology links are measured but are not required on every SOP, and the result is not a recommendation for boilerplate or sitewide insertion.
- Machine surfaces: the proposed build generates
/task-registry.json,/task-registry.schema.json,/llms.txt, and/evidence-summary.jsonfrom the same task data. Those paths remain local deployment artifacts until the pull request merges and a trusted Pages job passes the external governance gate. - Evidence trust fails closed: policy 1.0 has four empty maps:
trustedReceiptIssuers,trustedEvaluatorIssuers,trustedProductionDataIssuers, andtrustedOutcomeIssuers. E4 requires a trusted production-data signature over the complete event ledger and a separate trusted outcome-issuer signature over the downstream outcome. Those roles are not interchangeable. With all four maps empty and zero ledger records, no real claim can qualify above E0. - Source integrity is bounded: each resolved skill has an exact source digest captured before presentation trimming plus a separate embedded-Markdown digest. These digests detect content changes; they do not independently attest authorship, ownership, publication, or time. Historical records, policies, and schemas are versioned and intended to be append-only, but that is preventive only when external branch governance requires the checks before merge.
- Agent-language finding remains visible: after three high-risk task repairs, 236 local skills still contain persistence, memory, or model-specific design language. The proposed build labels that language unverified design intent, lowers readiness, and keeps it at E0 until a frozen stack passes accepted trials.
Public defects already repaired
- 10/10 stale task URLs resolve to their expected canonical destinations. Three prior Rank Math rules looked active but stored serialized JSON as the source and did not match a real request. Eight stale anchors in the legacy Task Library were replaced with direct canonical links.
- The legacy Task Library hero now contextually links the words “Content Factory” to the canonical page; the saved link passed cache-busted Chrome verification.
- The canonical public-figure task now moderates behavior rather than sentiment, protects honest criticism, and includes authority, acceptance, rollback, escalation, and evidence boundaries.
- The Local Service Spotlight AI Builder, BlitzMetrics student, and teacher pages no longer imply unsupported earnings, payback, free-program, placement, or partner-funding outcomes. They grade disclosed AI use on accepted outcomes and keep live actions reviewer-gated.
- The legacy library no longer says AI is categorically incapable of helping with relationships or that an AI-assisted book creates “instant authority.” Human ownership of consent, trust, accuracy, originality, usefulness, and consequences is explicit.
- Blanket distribution, unenforced “B+,” “kill it forever,” and “every agent” language was narrowed to selective native distribution, explicit reviewer acceptance, evidence-backed owner decisions, and the current resolvable registry.
- The Meta-Article Prompt now forbids invented token/cost tables when telemetry is unavailable, separates measured values from estimates, rejects automatic-training and guaranteed-search-ownership claims, and requires artifact-aware capability and URL-repair receipts.
- The Content Factory page treats output yield, ranking, conversion, durability, structured data, and case-study performance as conditional and measurable. Scheduled agents propose or route review-ready work under configured permission and publishing gates.
- The live “Content Factory is an evidence loop” section connects the 10-year plan to task contracts, complete receipts, E0–E4 promotion, error-driven repairs, and an outcome-based college rubric that rejects AI-detector grading as a proxy for learning.
What the red team caught
A plausible first validator still allowed unrelated experiment arms to inflate a result, mutable stack labels to masquerade as reproducible releases, self-declared preregistration, future-dated downstream outcomes, evaluator/operator overlap, secret text outside URLs, mutable policy history, direct-push bypass, and false staleness from a stripped terminal newline. The first article audit also stripped meaningful trailing slashes, used a bot-style user agent that triggered WAF false positives, and exposed a real knowledge-panel canonical mismatch. Each defect became a validation rule, regression test, canonical source repair, or explicit operational boundary.
RSA trust-key hardening: a malicious or accidentally provisioned exponent of 1 could turn an encoded digest into a forgeable “signature,” including across the receipts needed for an E4 record. Trust-key metadata now fails before activation unless it has the exact supported fields and algorithm, exponent exactly 65537, and a canonical no-leading-zero, odd, 2048–8192-bit modulus coprime to that exponent. Verification also requires the decoded signature to have the modulus width and satisfy signature < n. Malformed metadata is rejected, and a dedicated exponent-1 forged-E4 regression proves that such a policy cannot promote the record.
The remaining trust boundary is explicit. Git commits and artifact bytes can be digest-bound, but the offline validator does not independently attest a host timestamp, source ownership, or publication. RSA PKCS#1 v1.5 SHA-256 verification is implemented for preregistration and evaluation receipts, yet the corresponding trusted issuer maps are empty. E4 adds two independent signatures: the production event ledger must verify against trustedProductionDataIssuers, and the downstream outcome must verify against trustedOutcomeIssuers. Synthetic keys exercise code paths but establish no repository trust. Remote, private, digest-only, or privacy-blocked artifacts cannot qualify as byte-verified evidence.
The proposed workflow can run the bounded, SSRF-safe public-article auditor in observation mode only on trusted non-PR builds after the repository-governance gate. It cannot configure branch protection itself and is not active base-repository automation while PR #1 remains unmerged and undeployed. This continuation updates the documentation artifacts only; it does not publish or deploy repository changes.
| Local registry | 240 / 240 records generated without the private tracker feed |
| Documentation coverage | taskPage: 39 / 240 records, 18 unique pages, 201 unmapped; 18 / 18 unique mapped pages clean. definitiveArticle: 199 / 240 records, 17 unique URLs, 41 gaps; 17 / 17 unique URLs clean. |
| Contextual link graph | 39 mapped records; 27 unique pages fetched; 0 failures; library 9, methodology 0, task→article 34, article→task 25, complete loops 5 |
| Scoring boundary | Scores belong to tasks; categories show member-task medians, distributions, modes, and evidence coverage—not job-replacement probability |
| Real E1–E4 records | 0; all four policy 1.0 trust maps are empty; current local evidence remains E0 |
| Automated checks | 141 tests passing in the authoritative full patched-tree discovery |
| Delivery boundary | PR #1 is publicly reviewable proposed source, but it is not merged or deployed. Maintainer review, workflow approval, a trusted build, and an active no-bypass PR/review/build/force-push/deletion governance ruleset are still required. |

