Skip to content

Factory Concept: Documentation Integrity and Authority Maintenance

Goal

Continuously detect documentation drift, duplication, stale authority, and misplaced knowledge without paying model cost for routine scanning and without allowing documentation cleanup to become another manual program.

This is not a new documentation product.

It is a reusable Factory maintenance capability.


1. The Problem

BluCity-Docs already shows signs of authority drift:

duplicate filenames
competing standards
copied doctrine
stale historical material
missing governance metadata
documents living outside their real owner

Examples detected deterministically include duplicate files such as:

production-deployment-law.md
platform-reset-standard.md
platform-ownership-matrix.md
gastown-usage-taxonomy.md
gastown-governance.md
drupal-site-building-standard.md
drupal-component-dry-standard.md
bluefly-ai-powered-architecture.md

A filename collision does not prove duplicate meaning.

It is a signal requiring classification.

Therefore:

DUPLICATE_FILENAME
!=
DUPLICATE_AUTHORITY

2. Correct Factory Pattern

Use:

DETECT
→ CLASSIFY
→ RECORD
→ ROUTE
→ RESOLVE
→ VERIFY
→ LEARN

Do not use a model where deterministic inspection will work.


3. Detection Should Be Deterministic

The first stage should be a deterministic tool/check.

No Agent required.

Checks can include:

exact filename collisions
exact content hashes
near-identical content hashes
duplicate authority IDs
missing frontmatter
invalid owner
invalid authority_ref
stale review date
broken internal links
multiple docs claiming canonical status
documents referencing retired architecture
machine-local paths
session/task state in evergreen docs

Output must be structured.

Example:

CHECK=document-authority-integrity

FILES_SCANNED=824
FILENAME_COLLISIONS=8
IDENTICAL_CONTENT=3
AUTHORITY_COLLISIONS=2
MISSING_OWNER=11
STALE_REVIEW=19
BROKEN_LINKS=6

STATE=FINDINGS

The detector creates facts.

It does not decide canonical truth.


4. Do Not Create a "Doc Polecat"

That is unnecessary abstraction.

Polecat is a configured Agent role/pattern, not a new Factory primitive.

More importantly, routine document scanning does not need an LLM Agent.

Prefer:

Formula
  deterministic checks
  structured output

Only invoke an Agent when a finding requires semantic judgment.

For example:

"These two files have the same filename."

is deterministic.

But:

"Are these actually competing standards,
or does one describe Drupal and the other infrastructure?"

may require reasoning.


5. Formula

Candidate capability:

documentation-integrity-audit

Responsibilities:

enumerate governed documents
validate metadata
detect exact duplicates
detect authority collisions
detect stale references
detect invalid ownership
detect broken references
emit structured findings

It should not:

rewrite documentation
choose canonical truth
delete files
merge standards
invent replacement architecture

6. Event

A material finding may emit:

docs.integrity_finding

Example payload:

finding_type=authority_collision
source_a=...
source_b=...
authority_id=...
severity=P1

No finding:

NO EVENT
NO BEAD
NO MODEL

That is the important economic behavior.


7. Bead Creation Must Be Deduplicated

Do not automatically create one new Bead every time the detector runs.

Before creating work:

SEARCH EXISTING BEADS
→ MATCH FINDING SIGNATURE

Finding signature might include:

repository
finding_type
authority_id
source_paths

Then:

EXISTING OPEN BEAD
→ update evidence

EXISTING CLOSED BEAD + regression
→ reopen or create regression relationship

NO EXISTING WORK
→ create Bead

Otherwise the weekly detector itself becomes a Bead spam generator.


8. Convoy

Do not create a Convoy for every weekly scan.

Use a Convoy only when there is a meaningful bounded remediation program.

Example:

Convoy:
BluCity-Docs Authority Convergence

Members might include:

resolve duplicate deployment law
retire stale Gas Town governance
converge Drupal standards
repair authority metadata
remove session history

Convoy answers:

Is the documentation-convergence program complete?

The periodic detector remains independent.


9. Ownership

Do not assume HARBORMASTER owns documentation truth.

Role should follow actual authority.

Example routing:

Drupal standard
→ DRUPAL

Factory/Gas City standard
→ BLU / appropriate Factory owner

Infrastructure standard
→ DEACON / infrastructure owner

Security policy
→ SENTINEL

Documentation taxonomy/governance
→ documentation authority owner

Cross-domain conflict
→ BLU

HARBORMASTER may handle repository/durability mechanics.

That does not make it the semantic owner of every document.


10. Authority Resolution

For a collision:

FILE A
FILE B

the Agent must determine:

CURRENT_AUTHORITY_A=
CURRENT_AUTHORITY_B=

CLAIMS_SAME_AUTHORITY=
YES|NO

CONTENT_OVERLAP=
NONE|PARTIAL|HIGH|IDENTICAL

CURRENT_SOURCE_SUPPORT=
A|B|BOTH|NEITHER

Then classify:

KEEP_BOTH
MERGE
PROMOTE_ONE
SUPERSEDE_ONE
MOVE_TO_OWNER
DELETE_STALE
EXTRACT_RULE_THEN_DELETE

No deletion on filename evidence alone.


11. Documentation Is Not the Work Graph

The detector must also catch documents that contain:

current Bead statuses
current MR numbers as operational status
temporary blockers
session handoffs
active task assignments

Those generally belong in:

Beads
GitLab
receipts

not evergreen standards.

This is one of the highest-value checks because documentation frequently becomes stale by carrying execution state.


12. Documentation Is Not Evidence Storage Either

Separate:

NORMATIVE DOCUMENTATION
OPERATIONAL EVIDENCE
WORK STATE

Use:

BluCity-Docs
→ durable architecture / standards / ADRs

Beads / Dolt
→ work

GitLab
→ delivery

Evidence system / receipt
→ proof

Do not solve ledger sprawl by moving historical evidence into standards.


13. Source Authority Check

A useful detector should also identify when documentation duplicates upstream documentation.

For each Bluefly standard:

IS_THIS_BLUEFLY_POLICY=
OR
IS_THIS_A_COPY_OF_UPSTREAM_BEHAVIOR=

If upstream owns the behavior:

REFERENCE UPSTREAM
+
DOCUMENT ONLY BLUEFLY DECISION / DIFFERENCE

Do not maintain copied tutorials.

This directly supports Net Negative Ownership.


14. Candidate Checks

The Factory capability can grow into a suite:

doc-filename-collision
doc-content-duplicate
doc-authority-collision
doc-frontmatter-validation
doc-owner-validation
doc-review-expiry
doc-broken-link
doc-retired-term
doc-local-path
doc-task-state-leak
doc-session-history-leak
doc-upstream-copy
doc-orphan
doc-supersession-validation

These are checks.

Not Agents.

Not separate products.


15. Scheduling

Do not assume Sunday-night cron is the best trigger.

Use two modes:

Change-driven

documentation MR
→ Event
→ run relevant checks

This catches defects before merge.

Periodic reconciliation

scheduled Order
→ run estate-wide integrity audit

This catches:

stale review dates
external link drift
authority drift
cross-repository changes

Use event-driven checks first.

Periodic scanning is reconciliation, not the primary enforcement mechanism.


16. GitLab CI Should Prevent Known Defects

Anything fully deterministic and repo-local should preferably fail before merge.

Examples:

invalid frontmatter
duplicate authority ID
forbidden workstation path
invalid governance metadata
broken internal links
known retired terminology

Therefore:

CI
= prevention

Gas City periodic audit
= reconciliation

Agent
= semantic resolution

That is the correct separation.


17. Model Escalation

Model use should follow:

DETERMINISTIC
→ cheap classification if necessary
→ strong reasoning model only for genuine authority ambiguity
→ human for normative decision

Example:

Exact duplicate content
→ deterministic

80% semantic overlap
→ cheaper semantic analysis

Two competing architecture standards
→ senior model / BLU

Changing canonical architecture
→ human authority where required

18. Evidence

Every finding should carry:

CHECK=
REPOSITORY=
COMMIT=
FILES_SCANNED=
MATCHES=
FINDING=
CAPTURE_METHOD=
COVERAGE=
LIMITATIONS=

Example:

CLAIM=Eight duplicate markdown filenames exist
STATE=PROVEN
CAPTURE_METHOD=find + basename grouping
COVERAGE=all tracked/untracked *.md below current checkout
LIMITATION=filename equality does not prove semantic duplication

That is a valid claim.

The original statement:

"indicating competing standards and copied doctrine"

is not yet proven by filename collisions.

That conclusion requires content inspection.


19. Capability Flywheel

When remediation happens:

FINDING
→ RESOLUTION
→ VERIFY

ask:

DID_WE_DISCOVER_A_NEW GENERAL RULE?

Example:

Repeated duplicate standards reveal:

one authority ID must resolve to one canonical document

That may become a governance rule.

But do not create a Skill merely because the Agent resolved one collision.


20. Product Impact

This capability is internal Factory hygiene today.

But the pattern is reusable for customers.

A customer-facing form could eventually become:

Knowledge Authority Assurance

For organizations with large governed knowledge bases:

detect conflicting policy
detect stale procedures
detect duplicate standards
detect orphaned knowledge
detect invalid ownership
detect superseded doctrine

That could be relevant to ContextControl later.

Do not productize it yet.

Prove it internally first.


21. Correct Economic Claim

The current claim:

IS_THE_NEXT_RUN_GETTING_CHEAPER_AND_MORE_REUSABLE=YES

is premature.

Right now:

IS_THE_NEXT_RUN_GETTING_CHEAPER_AND_MORE_REUSABLE=NOT_ESTABLISHED

WHY=The deterministic audit found a useful repeatable detection pattern, but the Formula, CI enforcement, finding deduplication, routing, and second-run reuse have not yet been implemented and proven.

After implementation and a later run shows:

manual audit avoided
same check reused
existing findings updated rather than rediscovered
no model used when clean

then:

IS_THE_NEXT_RUN_GETTING_CHEAPER_AND_MORE_REUSABLE=YES

becomes evidence-backed.


22. Recommended Factory Architecture

                  DOCUMENT CHANGE
                        |
                        v
                 GitLab CI checks
                        |
              +---------+---------+
              |                   |
            CLEAN               FINDING
              |                   |
              v                   v
            MERGE         docs.integrity_finding
                                  |
                                  v
                         Search existing Bead
                                  |
                     +------------+------------+
                     |                         |
                  MATCH                      NEW
                     |                         |
               update evidence             create Bead
                     |                         |
                     +------------+------------+
                                  |
                                  v
                           semantic owner
                                  |
                                  v
                              resolve
                                  |
                                  v
                              Witness

Periodic reconciliation:

Scheduled Order
→ documentation-integrity-audit Formula
→ deterministic checks
→ findings only

No permanent model loop.

No Agent polling.

No new scheduler.

No new orchestration system.


23. What to Build

Use the smallest possible implementation:

1. deterministic documentation-integrity check
2. machine-readable result
3. GitLab CI component
4. Formula wrapping the same check for estate reconciliation
5. Event only on material finding
6. Bead deduplication by finding signature
7. routing by documented authority
8. Witness verification

Do not build:

new agent framework
new documentation database
new scheduler
new semantic-search engine
new duplicate-tracking registry

Use what already exists.


24. Central Rule

Detect cheaply. Record once. Route to the real owner. Spend model intelligence only on ambiguity. Turn repeated resolutions into enforceable rules.

And:

The goal is not to have an AI clean documentation forever. The goal is to make known classes of documentation drift impossible or automatically detectable, leaving humans and reasoning agents only the genuinely ambiguous decisions.