ArchiveBox/docs/ArchiveBox-Architecture-Diagrams.md

127 lines
5.8 KiB
Markdown

# ArchiveBox Architecture Diagrams
This page is a map of the current execution and persistence paths. The implementation lives primarily in:
- `archivebox/cli/` for CLI entry points
- `archivebox/services/runner.py` for crawl and snapshot execution
- `archivebox/crawls/models.py` for the `Crawl` model and its atomic queue transitions
- `archivebox/core/models.py` for `Snapshot`, `ArchiveResult`, and Snapshot queue transitions
- `archivebox/services/` for bus event projectors
- `abxpkg` and `abx-plugins` for binary resolution and plugin hooks
## High-Level Execution Flow
```mermaid
flowchart TD
ENTRY["CLI, Web UI, REST API, or scheduler"] --> CRAWL["Create or resume a Crawl row"]
CRAWL --> RUNNER["run_crawl() / CrawlRunner"]
RUNNER --> DISCOVER["Create or select Snapshot rows"]
DISCOVER --> EVENTS["Emit crawl and snapshot lifecycle events"]
EVENTS --> PLUGINS["Run selected abx-plugin hooks"]
PLUGINS --> PROCESSES["Persist Process rows and hook output"]
PROCESSES --> RESULTS["Project ArchiveResult rows"]
RESULTS --> FILES["Write snapshot output files"]
RESULTS --> SNAPSTATE["Seal or requeue Snapshot"]
SNAPSTATE --> CRAWLSTATE["Seal, pause, or continue Crawl"]
EVENTS --> BINREQ["BinaryRequestEvent"]
BINREQ --> ABXPKG["abxpkg resolution"]
ABXPKG --> HOST["Compatible host binary"]
ABXPKG --> MANAGED["Managed install fallback"]
HOST --> ENV["Project resolved binary into LIB_DIR/env/bin"]
MANAGED --> ENV
CRAWL -.-> DB["SQLite database"]
PROCESSES -.-> DB
RESULTS -.-> DB
FILES -.-> STORAGE["archive/users/... snapshot storage"]
```
ArchiveBox has one normal crawl execution path. CLI commands and web/API actions create or select database rows, then call the same runner. The runner emits lifecycle events, abx-plugin hooks do the extraction work, and service projectors persist processes and results.
Binary discovery and installation always goes through abxpkg. Compatible host binaries are preferred; managed providers are the fallback. Resolved binaries are projected into `LIB_DIR/env/bin` before programmatic use. `LIB_DIR/bin` is only a convenience directory for humans.
## Persistent Data
```mermaid
flowchart LR
DATA["ArchiveBox data directory"] --> DB["index.sqlite3"]
DATA --> ARCHIVE["archive/users/<user>/snapshots/<date>/<domain>/<uuid>/"]
DATA --> SOURCES["sources/"]
DATA --> LOGS["logs/"]
DATA --> LIB["lib/env/bin/ resolved binaries"]
ARCHIVE --> PLUGINOUT["Plugin-namespaced outputs"]
ARCHIVE --> META["Snapshot metadata and indexes"]
```
The database is the source of truth for model state. Snapshot directories contain captured artifacts and rendered metadata. Older collections may also contain legacy timestamp-named snapshot directories.
## `Crawl` Queue Lifecycle
Implemented directly by `Crawl` in `archivebox/crawls/models.py`. The database
row is the durable state; the runner claims `retry_at` with a conditional update
before it performs side effects, then calls the model's explicit lifecycle
methods. There is deliberately no second in-memory state machine that can drift
from the row owned by another process.
```mermaid
stateDiagram-v2
[*] --> QUEUED
QUEUED --> STARTED: runner claim and valid URLs
QUEUED --> QUEUED: claimed but not ready
QUEUED --> SEALED: all existing snapshots finished
STARTED --> SEALED: all snapshots finished
QUEUED --> PAUSED: pause requested
STARTED --> PAUSED: pause requested
PAUSED --> QUEUED: resume requested
PAUSED --> PAUSED: not runnable
QUEUED --> SEALED: explicit seal
STARTED --> SEALED: explicit seal
PAUSED --> SEALED: explicit seal
SEALED --> [*]
```
A crawl owns a set of snapshots. The runner creates or discovers those snapshots and projects crawl events while the row is `STARTED`; sealing waits for their normal lifecycle to finish. Pausing also schedules child snapshots to pause, and resuming returns the crawl to the runnable queue. Scheduled maintenance is dispatched directly by `CrawlSchedule`; it does not create a synthetic crawl or snapshot.
## `Snapshot` Queue Lifecycle
Implemented directly by `Snapshot` in `archivebox/core/models.py`, using the
same conditional `retry_at` claim protocol as `Crawl`.
```mermaid
stateDiagram-v2
[*] --> QUEUED
QUEUED --> STARTED: runner claim and URL is ready
QUEUED --> QUEUED: claimed but not ready
QUEUED --> SEALED: all existing results finished
STARTED --> SEALED: all hook results finished
QUEUED --> PAUSED: pause requested
STARTED --> PAUSED: pause requested
PAUSED --> QUEUED: resume requested
PAUSED --> PAUSED: not runnable
QUEUED --> SEALED: explicit seal
STARTED --> SEALED: explicit seal
PAUSED --> SEALED: explicit seal
SEALED --> [*]
```
The runner creates one queued `ArchiveResult` per selected hook, executes those hooks through the shared event bus, and seals the snapshot after every result reaches a final status. The narrow search-index maintenance operation on an already sealed snapshot is the intentional exception; it does not reopen or invent a second general lifecycle path.
## `ArchiveResult` Projection
`ArchiveResult` is not driven by a separate in-memory state machine. The runner creates queued rows, and `ArchiveResultService` projects `ArchiveResultEvent` and `ProcessCompletedEvent` data into them.
```mermaid
flowchart LR
QUEUED["queued"] --> STARTED["started"]
STARTED --> SUCCEEDED["succeeded"]
STARTED --> FAILED["failed"]
STARTED --> SKIPPED["skipped"]
STARTED --> NORESULTS["noresults"]
STARTED -. recoverable wait .-> BACKOFF["backoff"]
BACKOFF -. resumed work .-> STARTED
```
`succeeded`, `failed`, `skipped`, and `noresults` are final result statuses. Each row identifies the plugin and hook that produced it and stores structured output, file metadata, timing, and error details.