Struktur

extract

Main extraction command for Struktur CLI.

Synopsis

struktur [extract] [options]

extract is the default command — struktur --input file.pdf ... and struktur extract --input file.pdf ... are equivalent.

Input options (exactly one required)

Prop

Type

Schema options (exactly one required)

Prop

Type

--fields is the quickest way to define a schema without writing JSON. See --fields reference for the full syntax.

Model

Prop

Type

Supported providers: openai, anthropic, google, cerebras, opencode, openrouter, ollama.

OpenRouter routing

OpenRouter model names accept two kinds of suffix:

SuffixEffect
:nitroRoute to the fastest available provider (OpenRouter-native)
:floorRoute to the cheapest available provider (OpenRouter-native)
:freeRestrict to free-tier endpoints (OpenRouter-native)
#<provider>Pin to a specific upstream inference provider, e.g. #cerebras
# Pin to a specific upstream provider
--model "openrouter/anthropic/claude-3.5-sonnet#cerebras"

# Prefer the fastest provider — often roughly halves wall-clock time
--model "openrouter/deepseek/deepseek-v4.1-flash:nitro"

Routing variants are stripped before capability lookups such as vision support, so a variant resolves to its base model.

Parsing options

These flags control how --input files are parsed before extraction.

Prop

Type

Image options (PDF inputs)

Prop

Type

For custom screenshot dimensions, use struktur parse --screenshots --screenshot-scale <num> and pipe the artifact to struktur extract --artifact-file -.

Image overview

When --images is set, the PDF parser also composes a single image overview — a labelled contact sheet of all extracted images, with each thumbnail captioned by its virtual path. It is added to the artifact as one extra generated image.

The overview gives a vision model the whole visual context of a document for the cost of one image, so it can decide which images are worth viewing in detail instead of reading them all. It counts as a single image toward --max-images, and --prefill-images always considers it first.

Label thumbnails are scaled to a ~220px longest edge (120px floor) and the sheet is capped at a 1500px longest edge, so large documents produce several numbered overview sheets. Byte-identical images (recurring logos, letterheads) and images under 40px in either dimension are filtered out before compositing. Disable compositing with struktur parse --no-image-overview.

Strategy

Prop

Type

Strategy names: simple, parallel, sequential, parallelAutoMerge, sequentialAutoMerge, doublePass, doublePassAutoMerge, agent (default).

Agent strategy options

When using --strategy agent, these additional options are available:

OptionDescriptionDefault
--max-stepsMaximum tool calls per iteration50
--max-iterationsNumber of iteration loops1
--instructionsAdditional instructions appended to the strategy prompt—
--reasoning-effortReasoning effort for thinking models: low | medium | highmodel default
--prefillPre-load up to this many tokens of document text as synthetic read calls before the agent starts. Accepts 300k, 1.5m and 300_000.— (off)
--prefill-imagesMaximum images to pre-load alongside --prefill, image overviews first1
--images-outputWrite the extracted image map (virtual path → base64) to this file—

Context prefill

The agent normally spends its first steps reading /artifact.json and viewing the image overview. Prefill hands it that context up front, as completed read and view_image tool calls, so it can start extracting immediately:

# Pre-load up to 300k tokens of text and the image overview
struktur extract --input ./expose.pdf --schema ./schema.json \
  --prefill 300k

# Also load the three largest other images
struktur extract --input ./expose.pdf --schema ./schema.json \
  --prefill 300k --prefill-images 4

The prefill block is deterministic, so the same document always produces the same bytes and providers can prompt-cache the prefix. Prefilling text is cheap and safe; prefilling images is not — a single page image can be several megabytes of base64, which is why the default stops after the overview and the total payload is byte-capped.

Output

Prop

Type

Progress

When stderr is a TTY, a progress bar is shown:

◈ ▰▰▰▰▰▱▱▱▱▱ 50% | batch 2/5

The bar is suppressed in non-interactive mode (piped stderr).

To consume progress, tool calls, and reasoning programmatically instead, run with --format json and read the NDJSON events from stderr. See Events & Observability for the event contract.

Examples

echo "Invoice #1042 from Acme Corp. Total: $2,400.00." | \
  struktur --stdin -f "invoice_number, vendor, total:number" \
  --model openai/gpt-4o-mini
struktur --input invoice.pdf \
  --fields "invoice_number, vendor, total:number" \
  --model openai/gpt-4o-mini
struktur --input invoice.pdf --images \
  --schema invoice-schema.json \
  --model openai/gpt-4o
# Use parse for custom screenshot settings, then pipe to extract
struktur parse --input slides.pdf --screenshots --screenshot-scale 2 | \
  struktur --artifact-file - \
  --fields "title, slide_count:integer" \
  --model openai/gpt-4o
struktur --input report.txt \
  --schema-json '{"type":"object","properties":{"summary":{"type":"string"}},"required":["summary"],"additionalProperties":false}' \
  --model openai/gpt-4o-mini
cat document.md | struktur --stdin --schema schema.json --model anthropic/claude-3-5-haiku-20241022
struktur --input large.md --schema schema.json --model openai/gpt-4o \
  --strategy parallel --output result.json
struktur --input data.bin --mime application/pdf \
  --fields "title, author" --model openai/gpt-4o-mini
struktur --input report.docx --parser @myorg/docx-parser \
  --fields "title, summary" --model openai/gpt-4o-mini
struktur --input data.txt --schema https://myserver.com/schemas/invoice.json --model openai/gpt-4o-mini
struktur --input doc.pdf --fields "title" --model openai/gpt-4o-mini --debug

See also

On this page