Struktur

parse

Convert files to Artifact JSON for inspection or pre-processing.

Synopsis

struktur parse --input <file> [options]
struktur parse --stdin [options]

Description

Converts a file or stdin to Artifact JSON. Use this to:

Inspect

See how Struktur will represent your document before running extraction

Cache

Pre-process files and cache the artifact JSON for repeated extraction

Debug

Debug parser output when configuring a custom parser

Pipeline

Pipe artifacts into extract for decoupled workflows

Options

Input (exactly one required)

Prop

Type

Output

Prop

Type

Parser control

Prop

Type

PDF processors

--processor selects how PDF pages become text and images. External engines are loaded on demand, so the default keeps an install dependency-free.

ProcessorWhat it doesRequires
pdf-parseDefault. Fast per-page text extraction, no layout awareness.—
vlmRenders each page and asks a vision model to transcribe it. Best for scanned or heavily designed documents.Ghostscript (gs) and a model
doclingIBM Docling — layout analysis, tables preserved as markdown.pip install docling
liteparseLlamaIndex's lightweight layout parser.@llamaindex/liteparse npm package
kreuzbergKreuzberg's unified document extractor.@kreuzberg/node npm package

Image extraction (PDF inputs)

Prop

Type

When --images is set, the parser also composes an image overview: a contact sheet of the extracted images, each thumbnail captioned with its virtual path. It gives a vision model the whole visual context of the document for the cost of one image, so it can pick which images to inspect closely rather than reading them all. Thumbnails are scaled to a ~220px longest edge (120px floor) and sheets are capped at 1500px, so a document with many images produces several numbered sheets. Byte-identical repeats and images under 40px in either dimension are filtered out before compositing. Disable it with --no-image-overview.

Parser resolution order

  1. --parser <pkg> flag — bypasses all config
  2. Parser configured for the detected MIME type (struktur config parsers add ...)
  3. Built-in parser for the MIME type
  4. Error: no parser found — suggests struktur config parsers add

Built-in parsers

MIME typeBehavior
application/pdfPer-page text via pdf-parse (or the selected --processor). Add --images for embedded images, --screenshots for page renders. Add --images also for a labelled image overview.
text/*Split on double newlines into content slices.
image/*Single-content artifact with the image as a media item.
application/jsonIf it validates as SerializedArtifact[], passed through unchanged.

Examples

struktur parse --input document.pdf
struktur parse --input slides.pdf --images --screenshots --output artifact.json
struktur parse --input data.xlsx --parser @myorg/xlsx-parser
struktur parse --input doc.pdf --images | \
  struktur extract --artifact-file - --fields "title, author" --model openai/gpt-4o-mini
struktur parse --input doc.pdf | struktur utils artifact-viewer --stdin > viewer.html
open viewer.html

See also

On this page