Markdown Conversion
Vertesia Markdown Conversion transforms document pages into structured markdown that preserves reading order, headings, lists, tables, and figures. It is the text-preparation path for documents whose full narrative should remain available to search, agents, and prompts.
Markdown conversion is one part of Content Preparation. For transactional documents where structured fields matter more than the complete page narrative, use Intelligent Document Processing instead, or combine both in the same intake policy.
When to use it
Markdown conversion works well for:
- reports and regulatory filings
- contracts, policies, and manuals
- presentations and visually structured PDFs
- scanned documents that need OCR before conversion
- documents with tables whose relationships would be lost in a plain text dump
For simple digital files, the automatic conversion method can use a mechanical text path. Use the LLM conversion method when page layout and visual structure carry meaning.
Configure a content type
In Studio:
- Open Content > Types and select the content type.
- Open the Intake tab.
- Select Conversion.
- Enable text conversion and choose LLM as the method.
- Keep the output format set to Markdown.
- Save the policy.
You can also add conversion instructions, such as keeping only commercial terms, preserving all table notes, or excluding boilerplate appendices. Page ranges and the locate pass can limit conversion to the relevant pages of a long document.
The equivalent policy shape is:
{
"text_conversion": {
"enabled": true,
"method": "llm",
"output_format": "markdown",
"instructions": "Preserve headings, tables, figure captions, and footnotes."
}
}
See Intake Configuration for project defaults, per-type inheritance, page selection, and model-independent conversion options.
Process and inspect a document
When automatic intake is enabled, upload a document normally. Standard intake resolves its content type, applies the type policy, and stores the converted markdown as the object's text representation.
Open the object in Studio to inspect the prepared text. For PDFs, the processing result also provides page renditions and annotations that make it possible to compare the markdown with the source document.

The markdown is available anywhere object text is used, including search, embeddings, prompt templates, and agents.
Read the result with the SDK
After intake completes, fetch the object's prepared text:
const { text } = await client.objects.getObjectText(objectId);
console.log(text);
To configure the policy through the SDK:
await client.store.types.update(typeId, {
intake: {
text_conversion: {
enabled: true,
method: 'llm',
output_format: 'markdown',
},
},
});
When updating an existing type, include the other intake settings you want to preserve in the submitted policy.
Combine conversion with IDP
A content type can enable both markdown conversion and IDP. This is useful when downstream users need the complete document text as well as verified structured fields. If only the fields matter, disable page conversion and use an output rendering template to generate concise markdown from the extracted properties.
