parse
Convert files to Artifact JSON for inspection or pre-processing.
Synopsis
struktur parse --input <file> [options]
struktur parse --stdin [options]Description
Converts a file or stdin to Artifact JSON. Use this to:
Inspect
See how Struktur will represent your document before running extraction
Cache
Pre-process files and cache the artifact JSON for repeated extraction
Debug
Debug parser output when configuring a custom parser
Pipeline
Pipe artifacts into extract for decoupled workflows
Options
Input (exactly one required)
Prop
Type
Output
Prop
Type
Parser control
Prop
Type
PDF processors
--processor selects how PDF pages become text and images. External engines are loaded on demand, so the default keeps an install dependency-free.
| Processor | What it does | Requires |
|---|---|---|
pdf-parse | Default. Fast per-page text extraction, no layout awareness. | — |
vlm | Renders each page and asks a vision model to transcribe it. Best for scanned or heavily designed documents. | Ghostscript (gs) and a model |
docling | IBM Docling — layout analysis, tables preserved as markdown. | pip install docling |
liteparse | LlamaIndex's lightweight layout parser. | @llamaindex/liteparse npm package |
kreuzberg | Kreuzberg's unified document extractor. | @kreuzberg/node npm package |
Image extraction (PDF inputs)
Prop
Type
When --images is set, the parser also composes an image overview: a contact sheet of the extracted images, each thumbnail captioned with its virtual path. It gives a vision model the whole visual context of the document for the cost of one image, so it can pick which images to inspect closely rather than reading them all. Thumbnails are scaled to a ~220px longest edge (120px floor) and sheets are capped at 1500px, so a document with many images produces several numbered sheets. Byte-identical repeats and images under 40px in either dimension are filtered out before compositing. Disable it with --no-image-overview.
Parser resolution order
--parser <pkg>flag — bypasses all config- Parser configured for the detected MIME type (
struktur config parsers add ...) - Built-in parser for the MIME type
- Error: no parser found — suggests
struktur config parsers add
Built-in parsers
| MIME type | Behavior |
|---|---|
application/pdf | Per-page text via pdf-parse (or the selected --processor). Add --images for embedded images, --screenshots for page renders. Add --images also for a labelled image overview. |
text/* | Split on double newlines into content slices. |
image/* | Single-content artifact with the image as a media item. |
application/json | If it validates as SerializedArtifact[], passed through unchanged. |
Examples
struktur parse --input document.pdfstruktur parse --input slides.pdf --images --screenshots --output artifact.jsonstruktur parse --input data.xlsx --parser @myorg/xlsx-parserstruktur parse --input doc.pdf --images | \
struktur extract --artifact-file - --fields "title, author" --model openai/gpt-4o-ministruktur parse --input doc.pdf | struktur utils artifact-viewer --stdin > viewer.html
open viewer.htmlSee also
- config parsers — Configure custom parsers
- Document Parsing — Parser system overview
- Artifact Format — Output format
- utils artifact-viewer — Visualize parsed artifacts