> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.instabase.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.instabase.com/_mcp/server.

# Extracting data from documents

> Specify classes and fields that define your project schema.

To process documents, you must specify which data points, or *fields*, you want to extract. If your project includes different document types, like a mix of passports and driver’s licenses, create a *class* for each document type and specify fields for each class.

You can create up to 250 classes per project, and up to 250 fields per class.

The classes and fields in your project form the project *schema*: the blueprint of information you want to extract from documents.

> **Tip**
>
> There are several ways to [reuse schemas from existing projects](/automate/projects#creating-projects) to save time.

## Autogenerating a project schema

Agent mode

Projects in agent mode can use a model to autogenerate classes and fields based on uploaded documents.

Autogenerated schemas create up to 20 fields per class using basic field types like text extraction and list extraction. This approach works well for straightforward extraction needs or as a starting point for more complex schemas that you can refine.

> **Before you begin**
>
> You must have uploaded a set of files that represent the types of documents you want to process. Five or so files of each type is a good start.

1. In the editing panel, click **Autogenerate schema**.

   If your project includes multipage files, you're prompted to optionally enable splitting files, also called split classification.

   > **Tip**
   >
   > Configure file splitting based on your *production processing requirements*. For example, if you intend to process compiled PDFs, enable file splitting even if your project includes only separate files. You can manually enable or disable file splitting in [project settings](/automate/projects#project-settings).

   It might take several minutes for the schema to generate.

   You can use the schema as-is, modify it, or click the **Undo schema** icon ![Icon that looks like a U-turn to the left.](/_fern-img/dcd4756425b7fce6dd159cfa701e1dfe98ccbccbbb5505fc98b169bcbd5f24c0.webp) to clear the autogenerated schema and manually generate your project schema.

## Manually generating a project schema

To manually create your project schema, create classes for each document type in your project, if necessary, then create fields for the data points you want to extract. This approach gives you complete control over your schema.

Manual schema generation works across all processing modes and provides access to all field types.

> **Before you begin**
>
> You must have uploaded a set of files that represent the types of documents you want to process. Five or so files of each type is a good start.

### Creating classes

If your project includes different document types, start by creating a *class* for each document type. You can then specify a different set of fields for each class.

Organization members can import *prebuilt classes* from a library of established schemas, such as paystubs, invoices, bank statements, and utility bills.

In projects with classification, a default class called *other* is assigned to documents that can’t be classified. You can’t delete or modify this class.

1. In the editing panel, click the **Create classes** icon ![Icon that looks like a horizontal bookmark or tab, with a small plus in the lower right.](/_fern-img/11a241141f2d59987b242317a8751794f8f8a7e7bfd612ad704432fcf6400c4a.webp), then select one of these options based on your AI Hub subscription and project requirements:

   * **Create classes** -- Lets you create a custom class without any fields. If you select this option, enter a succinct name for your document type, then click **←** to exit the class editing panel.

   * **Browse prebuilt classes**  -- Lets you add common document types and their associated fields based on a library of available schemas. If you select this option, choose the prebuilt classes that you want to add and click **Add to project**.

2. Use the **Create classes** icon ![Icon that looks like a horizontal bookmark or tab, with a small plus in the lower right.](/_fern-img/11a241141f2d59987b242317a8751794f8f8a7e7bfd612ad704432fcf6400c4a.webp) to add more classes as needed.

3. When you’re done creating classes, click **Classify documents**.

   If your project includes multipage files, you're prompted to optionally enable splitting files, also called split classification. You can split by document--which lets the model determine where document breaks occur--or split each page--which creates a new document at every page break.

   > **Tip**
   >
   > Configure file splitting based on your *production processing requirements*. For example, if you intend to process compiled PDFs, enable file splitting even if your project includes only separate files. You can manually enable or disable file splitting in [project settings](/automate/projects#project-settings).

   Classes are assigned to your documents and documents are grouped by class in the document list. Any documents that can’t be classified are assigned the *other* class.

4. Verify classification. If documents weren't classified as expected, edit classes to improve your results.

   1. In a class that wasn't identified accurately, click the overflow icon ![Icon that looks like an ellipsis, with three horizontal dots.](/_fern-img/d914b39588e7d4497a792e5d1312834c4a1021f88d95b383971c064de6984918.webp), then select **Edit class**.

   2. Enter a description to help the model more accurately identify documents in the class, then click **←** to exit the class editing panel.

      > **Tip**
      >
      > Effective descriptions include unique identifying details about a document class.
      >
      > In projects that don't use file splitting, you can reference file extensions to help classify documents. For example, the description for an *Images* class might be *Files with image file extensions, like JPEG, PNG, and TIF*.
      >
      > As a best practice, limit class descriptions to 1,000 characters (4,000 maximum).

   3. Use the overflow icon to edit more classes as needed.

   4. When you're done editing classes, click **Classify documents**.

#### Classification function

Enterprise

Classification functions let you use a custom Python function to classify documents. When a classification function is enabled for a project, it replaces the default LLM classification. All documents are classified using the custom function logic instead of AI inference.

Custom classification functions can provide predictable, rule-based classification for documents that follow consistent patterns. As an alternative to relying on model inference, you can implement custom logic that identifies document types based on specific text markers.

> **Info**
>
> Single-tenant In agent mode, classification functions can call LLMs to enable classification through a combination of model inference and deterministic rules. This functionality is available in single-tenant deployments only. For details, see [Calling LLMs from custom functions](/automate/llms-custom-functions).

For example, you might use a classification function to classify insurance forms based on specific identifiers:

```python
def unnamed_custom_function(context, keys, class_names):
    """
    Classifies insurance documents based on form codes and titles.
    Returns the appropriate class, which must appear as a key in class_names.
    """
    document_text = context['document_text']

    # Check for form codes near the beginning of the document
    document_start = document_text[:2000].upper()

    if "RNST" in document_start or "REINSTATEMENT REQUEST" in document_start:
        return 'Reinstatement Form'
    elif "TXFT" in document_start or "TRANSFER OF POLICY" in document_start:
        return 'Transfer Form'
    else:
        return 'Other'
```

Classification functions accept these parameters:

| Parameter                  | Required? | Description                                                                                                                                                                          |
| -------------------------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `context`                  | Required  | Stores metadata about the document.                                                                                                                                                  |
| `context['document_text']` | Optional  | Retrieves the entire text of the document.                                                                                                                                           |
| `context['file_path']`     | Optional  | Retrieves the path to the uploaded file.                                                                                                                                             |
| `keys`                     | Optional  | Access custom variables and [organization secrets](/admin/secret-management). Use `keys['custom']['<key-name>']` for custom keys and `keys['secret']['<key-name>']` for secret keys. |
| `class_name_enum`          | Required  | Enum containing all available class names in the project. Use to ensure returned values match existing classes.                                                                      |

Classification functions must return a value from the `class_name_enum` that corresponds to an existing class in your project. If the function returns a class name that doesn't exist, the document is assigned to the default "other" class.

To add a classification function, add at least one class to your project schema. Then click **Edit classes**, click the overflow icon, and turn on the **Classification function** toggle.

For additional guidance about custom functions, see [Writing custom functions](/automate/custom-functions).

#### Splitting function

Enterprise

Splitting functions let you use a custom Python function to split multipage files into documents. When a splitting function is enabled for a project, it replaces the default LLM-based splitting. All files are split using the custom function logic instead of AI inference.

Custom splitting functions can provide predictable, rule-based splitting for files that follow consistent patterns. As an alternative to relying on model inference, you can implement custom logic that identifies document boundaries based on specific text markers, structural elements, visual objects, or a combination of criteria.

> **Info**
>
> Single-tenant In agent mode, splitting functions can call LLMs to enable splitting through a combination of model inference and deterministic rules. This functionality is available in single-tenant deployments only. For details, see [Calling LLMs from custom functions](/automate/llms-custom-functions).

Splitting functions process the entire file and return a list of dictionaries that specify page ranges and class labels for each document split. Each dictionary must include:

* `class_label` -- The class name for the document split
* `page_start` -- The starting page number (0-indexed)
* `page_end` -- The ending page number (0-indexed, inclusive)

For example, you might use a splitting function to split files based on barcode separator pages:

```python
def unnamed_custom_function(context, keys, class_names):
    """
    Splits documents based on barcode separator pages.
    Returns a JSON-serialized list of page range dictionaries.
    """
    import json

    page_by_page_text = context['document_page_by_page_text']
    splits = []
    current_start = 0

    for page_num, page_text in enumerate(page_by_page_text):
        # Check if this page is a barcode separator (mostly empty or contains barcode)
        is_separator = (
            len(page_text.strip()) < 50 or  # Mostly empty page
            'BARCODE' in page_text.upper()   # Contains barcode marker
        )

        if is_separator and current_start < page_num:
            # End previous document before separator
            splits.append({
                "class_label": "Document",
                "page_start": current_start,
                "page_end": page_num
            })
            current_start = page_num + 1  # Start next document after separator

    # Add final document if there are remaining pages
    if current_start < len(page_by_page_text):
        splits.append({
            "class_label": "Document",
            "page_start": current_start,
            "page_end": len(page_by_page_text) - 1
        })
    # If no splits were created (e.g., all pages are separators), create a single split for all pages
    elif not splits and page_by_page_text:
        splits.append({
            "class_label": "Document",
            "page_start": 0,
            "page_end": len(page_by_page_text) - 1
        })

    return splits
```

Splitting functions accept these parameters:

| Parameter                               | Required? | Description                                                                                                                                                                          |
| --------------------------------------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `context`                               | Required  | Stores metadata about the document.                                                                                                                                                  |
| `context['document_text']`              | Optional  | Retrieves the entire text of the file.                                                                                                                                               |
| `context['document_page_by_page_text']` | Optional  | Retrieves a list of text strings, one per page.                                                                                                                                      |
| `context['file_path']`                  | Optional  | Retrieves the path to the uploaded file.                                                                                                                                             |
| `keys`                                  | Optional  | Access custom variables and [organization secrets](/admin/secret-management). Use `keys['custom']['<key-name>']` for custom keys and `keys['secret']['<key-name>']` for secret keys. |
| `class_names`                           | Required  | List of available class names in the project. Use to ensure returned class labels match existing classes.                                                                            |

Splitting functions must return a JSON-serialized list of dictionaries. Each dictionary represents one document split and must include `class_label`, `page_start`, and `page_end` keys. All `class_label` values must match existing class names in your project. Page ranges must not overlap and must cover all pages in the file.

You can use splitting functions independently or together with classification functions. When used together, the splitting function determines page ranges and assigns initial class labels, then the classification function processes each split document to refine or override the class assignment.

To add a splitting function, open the **File splitting** tab of your [project settings](/automate/projects#project-settings), and turn on the **Split multipage files** toggle. For additional guidance about custom functions, see [Writing custom functions](/automate/custom-functions).

### Creating fields

Create fields for each of the data points you want to identify.

1. In the editing panel, click **Add field**.

2. Enter a field name or select a suggested field name, then press **Enter**.

   Data is extracted based on field name alone and the result is displayed.

3. Do one of the following, based on whether your result is accurate:

   * Accurate result -- Click **←** to exit the field editing panel and continue adding fields.

   * Inaccurate result -- [Edit the field](#editing-fields). When you're done editing, click **←** to exit the field editing panel and continue adding fields.

#### Editing fields

If a field doesn't return the results you expect, you have options to fine-tune the results.

Access the field editor for an existing field by hovering over the field and clicking the edit icon ![Pencil icon.](/_fern-img/b4ea19b15837e4b064b8244bc307d55b8513ecb0b9135f785e601917e4736f1b.webp).

Ways to fine-tune or improve your results include:

* Providing a more detailed field description or prompt describing the information you're looking for.

  > **Tip**
  >
  > As a best practice, keep field and attribute names under 48 characters and use a description or prompt for longer content up to 1,000 characters (4,000 maximum).

* Enabling field-specific settings, such as turning on **Long table extraction** for list extraction and table extraction fields. This setting, available in agent mode, delivers improved extraction results for multipage tables and lists in PDFs, and for large CSV and Excel tables. On spreadsheets, each enabled field gets its own dedicated extraction request per chunk. If **Spreadsheet preprocessing** is off in [project settings](/automate/projects#processing), preprocessing is auto-enabled for that field only.

* For legacy mode projects, you can try choosing a [different model](/overview/models/) in most field types. While the standard model offers faster processing, it's most effective on shorter documents. The advanced model performs better with longer documents, challenging formatting, and on fields using multistep reasoning and complex math.

* You might also try adjusting your project's [file processing or digitization settings](/automate/projects/), which can improve results for objects such as tables and checkboxes.

> **Info**
>
> For more best practices and other tips, see [Prompting guidance](/automate/prompts).

When you're done editing a field, click **Run** to see results and further refine your edits if needed.

## Field types

Choose the field type appropriate for the data you want to identify.

| Field type                                 | Enterprise lite+ | Used to...                                                                                                               |
| ------------------------------------------ | :--------------: | ------------------------------------------------------------------------------------------------------------------------ |
| Text extraction                            |                  | Extract strings of text or numbers, such as address, account balance, or filing status.                                  |
| Table extraction                           |         ✓        | Extract structured tabular data from documents.                                                                          |
| List extraction                            |         ✓        | Extract multiple similar items with optional attributes, such as transactions or line items.                             |
| Document reasoning                         |                  | Generate results that aren't explicitly stated, through deduction, summarization, or calculation.                        |
| Visual reasoning                           |         ✓        | Analyze visual and stylistic elements including images, watermarks, layout, and formatting. Requires the advanced model. |
| Derived                                    |         ✓        | Generate values based on other fields in the class.                                                                      |
| [Custom function](#custom-function-fields) |         ✓        | Compute values or import external data using Python functions.                                                           |

For more guidance, see [Choosing field types](/automate/prompts#choosing-field-types).

### Custom function fields

The custom function field type lets you use a Python function to compute values or import data to your project schema.

You can write custom function code directly in the field editor, or you can import a [shared function](/automate/using-shared-functions) from the function library. Shared functions enable reuse of common functions across multiple fields and projects, making it easier to maintain and update your code.

For example, you might use a custom function to calculate total invoice amount using existing subtotal and tax rate fields:

```python
subtotal = float(subtotal)
tax_rate = float(tax_rate) / 100
tax_amount = subtotal * tax_rate
total_amount = subtotal + tax_amount

return round(total_amount, 2)
```

Custom function fields accept these parameters:

| Parameter                  | Required? | Description                                                                                                                                                                          |
| -------------------------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `context`                  | Required  | Stores metadata about the document.                                                                                                                                                  |
| `context['document_text']` | Optional  | Retrieves the entire text of the document.                                                                                                                                           |
| `context['file_path']`     | Optional  | Retrieves the path to the uploaded file.                                                                                                                                             |
| `keys`                     | Optional  | Access custom variables and [organization secrets](/admin/secret-management). Use `keys['custom']['<key-name>']` for custom keys and `keys['secret']['<key-name>']` for secret keys. |
| `<additional-field-name>`  | Optional  | When writing custom functions in automation projects, click **Add argument** to select additional fields in the class to use in the function.                                        |

#### Return type

When defining a cross-class custom function field, you can set the *return type* to **Text choices** to limit valid returned values to a specified set. In the custom function field editor, use the return type dropdown to switch from **Text** to **Text choices**, then add the allowed values. Define as a comma-separated list or use **Import as CSV** to upload a CSV file with one value per cell (up to 1,000 options).

The custom function must return one of those values, otherwise a validation error is shown. In [human review](/automate/review), fields with text choices display a dropdown of valid options from which reviewers can select.

> **Info**
>
> For additional guidance about creating and using custom functions, see [Writing custom functions](/automate/custom-functions).

## Viewing results across documents

To quickly scan or compare results, click the **Results table** icon ![Icon that looks like a 2x6 table.](/_fern-img/c9e926338f503d1a004df6d3b598b531d29ecac5c74cbb43985c16f5c5ec8d0c.webp) in the **Documents** header.

The results table corresponds to the current view in the editing panel, so the results you see change depending on your current task.

| If the editing panel shows...                           | Then the results table displays...                                                              |
| ------------------------------------------------------- | ----------------------------------------------------------------------------------------------- |
| Classes                                                 | Final results for all fields, across all classes.                                               |
| Field editor                                            | Final result and, if applicable, confidence threshold validation result for the selected field. |
| Validations with no rule selected                       | Validation results for all fields, across all classes.                                          |
| Validations with a rule selected *or* Validation editor | Validation result for the selected rule and the result of any fields used to calculate it.      |

## Reordering fields

To change the order of fields in the field editor, use the up and down arrows that display when you hover over a field.

Reordering fields can be necessary when creating [derived fields](#editing-fields), which can reference fields that precede it in the field editor. Additionally, reordering fields can be helpful to speed up reviews or support downstream integrations, because fields are displayed in processed results in the same order as in the field editor.

> **Warning**
>
> If you have derived fields or custom functions in your project that reference preceding fields, be aware that reordering fields can break the reference.

## Hiding fields

Hiding intermediate or computational fields can help simplify human review and downstream integration output.

Consider hiding fields that are used exclusively as input for derived fields or custom functions. For example, you might extract individual date components in separate hidden fields, then combine them into a final formatted date field that reviewers and downstream systems actually need.

To mark a field as hidden, open the field editor and enable **Hide field**.

Hidden fields can't have validation rules, because validations on hidden fields could create confusing review scenarios. If you hide a field with an active validation rule, the rule is removed. If you later unhide the same field, any previous validation rules are restored.

Hidden fields use processing resources and count toward field limits, but their visibility varies across different AI Hub interfaces:

| Interface                         | Hidden field behavior                                              |
| --------------------------------- | ------------------------------------------------------------------ |
| App run results (UI)              | Hidden by default, can be unhidden with human review field filters |
| App run results (exported)        | Unhidden                                                           |
| Accuracy tests                    | Hidden by default, can be unhidden via test configuration          |
| Human review                      | Hidden by default, can be unhidden with human review field filters |
| Deployment run results (exported) | Unhidden                                                           |
| Downstream integrations           | Hidden by default, can be unhidden via deployment configuration    |
| API & SDK results                 | Hidden by default, can be unhidden via deployment configuration    |
| Deployment metrics                | Hidden by default, can be unhidden via deployment configuration    |