Skip to main content

Documents

A document in architxt is a source artifact — a PDF, Markdown file, image scan, or plain text — that you upload so architxt can read it, extract structure from it, and connect it to the contextual graph.

Documents are the primary inputs into the system. Everything else — entities, relationships, mental models, evidence, and Hindsight memory bank sync — starts from a document.

This page describes the document's contextual role, the data you see and edit in the UI, and the lifecycle that moves a document from upload to extracted source.

What a document carries​

Conceptually, a document is a bundle of three things:

LayerWhat it holdsWhy it matters
SourceThe original file, its filename, path, and an optional External ID.Provenance. You can always re-download or re-process the source.
MetadataContext, tags, authors, and a user-defined Document Date.Organization and filtering across the document corpus.
Extracted contentMarkdown content, content hash, extracted images, and structured blocks.The material the graph builders, queries, and Reflect features operate on.

Source​

  • Filename is the original uploaded file name.
  • Full Path is the stored location, which may be a local path or URL.
  • External ID is a stable identifier you can supply during upload or edit later in the Document Details dialog.
  • Generated By records whether the file came from a user upload (user) or an import process (import).

Metadata​

A document can be linked to:

  • One Context — a bucket such as "Architecture" or "Research 2026". You set this from the documents table using the Context button or inside Document Details.
  • Many Tags — ad-hoc labels used for grouping and bulk operations. The Tags button opens a dialog to add or remove tags for selected documents.
  • Many Metadata entries — key/value pairs such as author or revision. The Metadata button edits these in bulk.
  • Authors — a list of names parsed from the file. You can add or remove authors in the Document Details dialog.
  • Document Date — the document's own date, independent of when it was uploaded. Set this via the Config bulk action or in Document Details.

Extracted content​

After the extraction pipeline runs, architxt stores:

  • Document Content — a Markdown representation of the document.
  • A content hash — used to detect changes and verify integrity.
  • Content blocks — optional structured blocks produced by smart editing.
  • Extracted images are written to disk next to the source file.

You view this on the Document Content tab inside Document Details, where you can also switch between highlighted entity tags and plain text.

How documents fit into architxt​

Upload
│
▼
┌──────────────┐
│ Document │ source + metadata
└──────┬───────┘
│ Extraction
▼
┌──────────────┐
│ Content │ markdown + images
└──────┬───────┘
│
┌─────┴─────┐
▼ ▼
Entities Mental Models
│ │
└─────┬─────┘
▼
Contextual Graph
│
▼
Reflect /
Hindsight Sync

Documents are not themselves graph nodes. They are the evidence behind nodes. Entity and edge records in the contextual graph can trace back to the document (and even the chunk) that supported them, which makes answers in Reflect grounded and verifiable.

Smart Edit​

After extraction, a document's content can be refined with Smart Edit. This is not a plain text editor: it breaks the Markdown into typed blocks — headings, paragraphs, images, fenced code, and tables — so you can work at the level of document structure rather than raw characters.

You reach Smart Edit from the Document Details dialog. The editor shows a structural sidebar and a content pane. From there you can:

ActionWhat it does
Remove or restore a heading sectionMarks every block under that heading as removed, or brings them back.
Remove or restore a single blockToggles deletion for an image, code block, table, or paragraph.
Edit a blockChanges the Markdown text of an individual block while preserving the rest.
Hide or show removed blocksToggles visibility of deleted material so you can review what was cut.
DiscardAbandon all changes and close the editor.
SaveWrites the reconstructed Markdown and the updated block list back to the document.

Smart Edit is useful for cleaning up extraction artifacts: removing boilerplate pages, fixing a misrecognized table, or deleting appendix sections before the content is promoted into the contextual graph.

Document table​

The documents table in the architxt UI surfaces the most important facts about each document:

UI labelWhat it showsWhy it matters
Document IDInternal architxt identifier.Used throughout the UI and API to refer to the document.
External IDStable identifier you supply during upload or import.Maps to Hindsight's document_id, so the same id can survive re-imports.
StatusCurrent lifecycle state such as Uploaded, Extracting, or Extracted.Tells you whether the document is ready to use or still being processed.
Document ContentExtracted Markdown representation of the file.The material the contextual graph, mental models, and Reflect operate on.
Content Hashsha256 fingerprint of the extracted content.Lets architxt detect when content has changed and avoid duplicate work.
Full PathStored location of the original file.Used for re-download, re-processing, and integrity checks.
Processing HistoryLog of every status transition.Audit trail you can inspect in Document Details.
ProgressOptional extraction progress object.Shown while a document is actively being extracted.
ContextThe document's assigned context, if any.Groups related documents and is sent to Hindsight during sync.
Document DateTimestamp you assign to the document itself.Independent of upload time; useful for filtering and sorting.

Document states​

The lifecycle has five primary states and one transient request state. The UI labels used in the documents table and Document Details dialog are shown alongside the internal status names:

Internal statusUI labelMeaning
uploadedUploadedFile is stored but extraction has not started.
ready_to_extractExtracting (queued)Queued for processing; a daemon can claim it.
processing_extractExtracting (active)Locked by a daemon and actively being extracted.
processed_extract_successExtractedExtraction finished and the Markdown content was saved.
processed_extract_failedExtracted — FailedExtraction failed; details are in Processing History.
request_release (transient)ReleasingUser asked to cancel while the document was extracting.

The documents table uses filter buttons named All, Uploaded, Extracting, and Extracted. Extracting covers both ready_to_extract and processing_extract. Extracted covers both success and failure states.

Lifecycle flow​

uploaded ──► ready_to_extract ──► processing_extract ──► processed_extract_success
│ │
└────────────────────────► processed_extract_failed

Upload​

When a file is uploaded through POST /documents (server/src/routes/documents-core.js), the server:

  1. Creates a database record with status uploaded.
  2. Writes the file to the configured storage path.
  3. Updates the stored Full Path with the final location.

If file storage fails, the database record is rolled back.

Queue for extraction​

Clicking Extract on selected documents transitions them to ready_to_extract. The route only allows this from one of these source states:

  • uploaded
  • processed_extract_failed
  • processed_extract_success

The transition is atomic: the status is updated only if the document is in an allowed state. If it is not, the route returns a STATE_CONFLICT error. The Extract button in the UI dynamically labels itself, for example Extract (new 2, reprocess 1), depending on how many selected documents are new versus already extracted.

Claim and process​

The extraction daemon (server/src/daemons/extract-daemon.js) polls for the oldest ready_to_extract document, then attempts to claim it. The claim is also atomic: a document is updated to processing_extract only if it is still ready_to_extract. If another daemon claimed it first, the loser drops it and moves on.

Once claimed, the daemon:

  1. Fetches the original source file.
  2. Runs the extraction pipeline.
  3. Reports the result back.

Report success or failure​

The extracted result is reported with a success flag and either Markdown (when successful) or an error (when not). The route:

  • Computes a sha256 hash for successful extractions.
  • Stores the content, hash, and any extracted images.
  • Writes a Processing History entry.
  • Sets the final state to processed_extract_success or processed_extract_failed.

Only documents currently in processing_extract can accept an extracted result.

Cancel and release​

While a document is Extracting, you can cancel it. The UI sets the status to request_release. The daemon checks for this status during processing. If it sees request_release, it releases the document back to uploaded.

Release has one extra behavior: if the document already has content and the sha256 hash matches, the release preserves that content and lands in processed_extract_success instead of uploaded. This prevents throwing away useful work when a cancel request arrives just after extraction finishes.

Orphan recovery​

If a daemon crashes while a document is processing_extract, the next daemon startup or periodic reconciliation detects orphaned documents. Documents stuck in Extracting for longer than the configured orphan_threshold_minutes are reset to Uploaded so they can be claimed again.

State transitions at a glance​

FromOperationTo
uploadedClick Extractready_to_extract
processed_extract_failedClick Extractready_to_extract
processed_extract_successClick Extractready_to_extract
ready_to_extractDaemon claims workprocessing_extract
processing_extractExtraction succeedsprocessed_extract_success
processing_extractExtraction failsprocessed_extract_failed
processing_extractUser cancelsrequest_release
request_releaseDaemon releasesuploaded or processed_extract_success
processing_extractDaemon releasesuploaded or processed_extract_success
processing_extractOrphan reconciliationuploaded

Processing History​

Every state transition is recorded in Processing History as a JSON entry. You can view this history on the Processing History tab inside Document Details. Each entry has this shape:

{
"timestamp": "2026-09-16T12:34:56.789Z",
"from": "uploaded",
"to": "ready_to_extract",
"success": true,
"reason": "User queued for processing"
}

Failure entries also include an error message and metrics. Helper functions in server/src/utils/db-helpers.js keep the format consistent across routes.

What you can do after extraction​

Once a document reaches Extracted:

  • Its content is indexed for full-text search.
  • Entities and relationships derived from it can be promoted into the contextual graph.
  • It can be re-queued for extraction if the content changes or extraction settings are updated.
  • It can be synced to a Hindsight memory bank.

Documents remain in the table unless explicitly deleted; architxt does not use a deleted flag. The Processing History preserves the full lifecycle for audit purposes.