Vertesia Documentation

Concepts

This page maps the Vertesia platform: what each building block is, and where to read more. Each section of the documentation covers one of these areas in depth.

The building blocks

ConceptWhat it isLearn more
PromptsReusable text templates with variable injection, assembled into the prompt for a task.Prompts
InteractionsA task you ask a model to perform: prompts, an optional result schema, and a model configuration.Interactions
AgentsAn interaction that runs as an autonomous loop, calling tools and reasoning over their results until it reaches an answer.Agents
ProcessesDurable state-machine workflows with named steps, deterministic routing, human pauses, and child subprocesses.Processes
Content Types & CollectionsStructured document storage: JSON Schema validation for document shape, collections for organization.Intake Configuration
Content PreparationTurning uploaded documents into data your agents can use — extraction, markdown conversion, embeddings.Content Preparation
ViewsReusable content navigation, faceted or agentic search, and result layouts, rendered in Studio or your own app.View Experiences
EventsPlatform events and the subscriptions that route them to a workflow, agent, process, or webhook.Event Subscriptions
EnvironmentsConnections to the LLM inference providers that run your models.Environments
Data PlatformSQL data stores and dashboards for structured data and analytics.Data Platform
Access ControlRoles and RBAC grants for platform actions; ABAC rules for per-document access.Access Control
Custom ApplicationsApps, MCP servers, and UI plugins that extend Studio or embed it elsewhere.Custom Applications

The Studio Assistant can build most of the above for you, and is the fastest way to get a first version of something working.

Choosing building blocks

Two decisions come up early, and getting them right saves the most rework.

Direct call or an agent?

An interaction always has a prompt and usually a schema. The question is whether the model needs to act or just respond.

Use direct inference — a single model call — when everything the model needs is already in the input. Classify, extract, summarize, translate, rewrite. It is faster, cheaper, and predictable.

Use an agent when the task needs tool use or iteration:

  • it must look things up — search, fetch documents it cannot name in advance, call an external API;
  • each step depends on the previous one;
  • the reasoning is too hard to land reliably in one pass and benefits from planning and revision;
  • a person needs to steer it mid-run.

Start with direct inference and promote to an agent when the task clearly needs it. "It might occasionally need to search" is not a reason — that is just a larger prompt.

An agent or a process?

Use an agent when one autonomous worker can reason its way to the result.

Use a process when the workflow itself is part of what you are delivering and needs explicit named steps, deterministic routing, durable pauses for human work, reusable subprocesses, or fixed parallel branches.

Business workflows of the kind you would normally draw as a flowchart map to processes, not agents.

Agent capabilities

An agent's abilities come from the tools it can call. Tools are granted through skills — a skill bundles a whole domain's tools together with instructions on how to use them, so you grant one skill rather than a list of individual tools.

For example, a research agent needs web search and document search; a document-authoring agent needs the authoring and file-editing skills. See Skills.

A single agent run can also fan out into parallel workstreams, each a child run with its own task.

Runs

A run is one execution of an interaction call — both the request to the model and the response from it.

  • Name
    created
    Description

    The run exists but has not started. Typically waiting for a client to begin streaming.

  • Name
    processing
    Description
    The run is executing.
  • Name
    completed
    Description
    The run finished successfully.
  • Name
    failed
    Description

    The run failed. The reason is in the error field.

Prompt template syntax

Each prompt template has a content_type that determines how variables are injected into its content.

  • Name
    Handlebars (recommended)
    Description

    Handlebars syntax with double curly braces. Supports {{variable}} references, conditionals ({{#if}}/{{else}}/{{/if}}), loops ({{#each items}}), and helpers such as {{_now}} and {{stringify obj}}. Set content_type to "handlebars".

  • Name
    JS Template (advanced)
    Description

    A JavaScript template engine in a sandbox, for composition beyond what Handlebars can express. Standard interpolation (${var}), control blocks, and array functions, plus a _ utility object with _.loadCsv(), _.jsonToCsv(), _.stringify(), _.addLineNumbers() and _.dayjs(). The template must return a string. Set content_type to "jst". Use it for data transformations, CSV processing, or programmatic prompt construction.

Inference providers

Environments connect to the platforms that run generative models. Vertesia supports:

Virtual providers assemble several models into one synthetic model with a balancing or mediation strategy:

  • virtual_lb - load balancing and failover across multiple models
  • virtual_mediator - multi-head execution with LLM mediation

See Environments for provider setup.

Language model background

If generative AI is new to you, these are the ideas the rest of the documentation assumes.

Large language models interpret and produce human language — including programming languages. They are pre-trained on very large content sets and can be fine-tuned afterwards for specific knowledge, though Retrieval Augmented Generation (RAG) is usually the better way to give a model access to your own data. Chatbots were the first well-known application, but a model is more usefully thought of as a processing entity you can call automatically, which is what opens up the range of use cases this platform is built for.

TermDescription
Context WindowHow much text the model can consider at once, measured in tokens. Like a notebook with a fixed number of pages: once it is full, only what is written there can be seen. Each model has its own limit.
TokenThe unit of text a model processes — a word, part of a word, a character, or punctuation. Splitting text into tokens is called tokenization.
Max tokensThe maximum number of tokens the model may generate in its response. On some models this counts against the context window.
EmbeddingsHigh-dimensional vectors representing text in a way that captures meaning, learned during training. They are what makes semantic search possible.
SimilarityHow close two pieces of text are in meaning. "Happy" and "joyful" sit near each other in embedding space; "happy" and "fast" sit far apart.

Was this page helpful?