> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.instabase.com/flow/guides/packet-processing-flows/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.instabase.com/_mcp/server. # Packet-processing flows > Design flows for packet processing. Enterprise Single-tenant This guide covers how to manually build packet-processing flows when your automation project defines cross-class fields -- that is, packet-level fields consolidated across document classes. To run the flow and inspect cross-class field output end to end, [publish your flow as an advanced app](/flow/advanced-apps). For schema and validation concepts in the app editor view, see [Extracting data from packets](/automate/packet-schema) and [Validating packets](/automate/validating-packets). ## Step order A typical packet-processing pipeline runs in the following step order: 1. *Process files* -- OCR and digitization. 2. *Unified extractor* -- Structured extraction across the packet using your configured schema. 3. *Apply checkpoint* -- Class-level validations (when your project defines them). 4. *Combine* -- Merges all file records into one stream before packet-level refinement. See [Combine classes](#combine-classes). 5. *Process case* -- Runs the packet refiner program (`case.ibrefiner`) that materializes cross-class fields into the IB stream. See [Process case](#process-case). 6. *Apply checkpoint* (optional) -- Cross-class validations, only when the published project defines rules on cross-class fields. See [Cross-class checkpoint](#cross-class-checkpoint). 7. *Combine* -- Final merge when you need all classes and cross-class output in one stream. See [Combine all](#combine-all). ## Combine classes Use the *Combine* step to present one combined stream to the *Process case* step. | Setting | Value | | --------------------------- | ----- | | `combine_same_file_records` | `no` | | `combine_all_file_records` | `yes` | ## Process case The *Process case* step loads a packet [refiner program](/flow/step-config-reference/creating-refiner-programs) from the flow bundle and writes output that includes materialized cross-class fields. For UI field descriptions, folder keys, `kwargs`, and the full reference for authoring `case.ibrefiner`, see [Process case](/flow/step-config-reference/process-case). | Kwarg | Purpose | | ------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `input_folder` | Combined packet IB input after class-level extraction -- typically the output of the preceding *Combine* step. | | `output_folder` | Output of the packet step, where cross-class fields are merged in. | | `case_program_path` | Path to the packet refiner program in the flow bundle; typically `modules/case.refiner/prog/case.ibrefiner`. When creating the module manually through the flow interface, you must create the `prog` folder yourself and move your `.ibrefiner` file into it. | | `settings` | Optional; often `{}`. | ```json { "kwargs": { "input_folder": "s4_combine", "output_folder": "s5_case", "case_program_path": "modules/case.refiner/prog/case.ibrefiner", "settings": {} }, "option": { "internalName": "process_case", "name": "Process Case" } } ``` ### Module layout The packet refiner program lives in the flow bundle at: ``` modules/case.refiner/ modules/case.refiner/prog/case.ibrefiner ``` *Process case* resolves `case_program_path` to that JSON program. The file does not have to be named `case.ibrefiner` -- the name is determined by the path you set in `case_program_path`. Per-class refiner modules and `.ibvalidations` files for class-level checkpoints live in separate module folders. > **Note** > > When you create a refiner module through the flow interface, the default file path is `modules/.refiner/.ibrefiner`. To match the expected path, use the file explorer to create a `prog` folder inside the `.refiner` folder, then move your `.ibrefiner` file into it. If you create your schema in a project, then [export the project to the flow editor](/automate/projects#exporting-and-importing-projects), the `prog` folder is created for you. ### Authoring packet refiner programs The fixed prefix builds the inputs for packet-level logic: a class-field manifest and one bridge column per referenced `(class name, field name)` pair. Bridge labels (typically `__ClassName_FieldName_`) are what downstream cross-class formulas reference. The cross-class field columns follow the bridges, one object per cross-class field in project order. Each uses a `choice_type`: `DERIVED`, `AGGREGATED`, or `UDF`. For schema details and native function descriptions, see [Process case](/flow/step-config-reference/process-case#authoring-packet-refiner-programs). The JSON examples below are drawn from a loan application project with four document classes: **Bank\_Statement**, **Loan\_Application**, **Pay\_Stub**, and **W2**. Substitute your real class and field names when authoring your program. #### Fixed-prefix skeleton Two patterns exist for the fixed prefix, depending on how your program was generated. **Pattern A: `get_intra_class_field` per bridge (common)** Each bridge column calls `get_intra_class_field` directly and includes its own `field_selection_strategy`. There is no separate shared aggregation column. This is the pattern generated when starting from an automation project. ```json { "dev_input": {}, "options": { "provenance_tracking": true, "auto_provenance": false }, "fields": [ { "label": "__class_field_inputs", "lines": [ { "function_id": { "name": "get_input_fields_per_class_field", "source": "NATIVE" }, "inputs": [ { "arg_name": "class_fields", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "ANY", "data_type_options": ["TEXT"], "value": "{}" } ] } ] }, { "label": "__Loan__Application_Loan__Amount_", "lines": [ { "function_id": { "name": "get_intra_class_field", "source": "NATIVE" }, "inputs": [ { "arg_name": "class_field_inputs", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "FIELD", "data_type_options": ["TEXT"], "value": "__class_field_inputs" }, { "arg_name": "class_name", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "ANY", "data_type_options": ["TEXT"], "value": "\"Loan_Application\"" }, { "arg_name": "field_name", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "ANY", "data_type_options": ["TEXT"], "value": "\"Loan__Amount\"" }, { "arg_name": "field_selection_strategy", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "ANY", "data_type_options": ["TEXT"], "value": "\"best_field_value\"" }, { "arg_name": "kwargs", "arg_type": "VAR_KEYWORD", "data_type": "DICT", "data_type_options": ["DICT"], "value": "{}" } ], "output_type": "TEXT" } ] }, { "label": "__W2_Wages__Tips_", "lines": [ { "function_id": { "name": "get_intra_class_field", "source": "NATIVE" }, "inputs": [ { "arg_name": "class_field_inputs", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "FIELD", "data_type_options": ["TEXT"], "value": "__class_field_inputs" }, { "arg_name": "class_name", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "ANY", "data_type_options": ["TEXT"], "value": "\"W2\"" }, { "arg_name": "field_name", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "ANY", "data_type_options": ["TEXT"], "value": "\"Wages__Tips\"" }, { "arg_name": "field_selection_strategy", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "ANY", "data_type_options": ["TEXT"], "value": "\"best_field_value\"" }, { "arg_name": "kwargs", "arg_type": "VAR_KEYWORD", "data_type": "DICT", "data_type_options": ["DICT"], "value": "{}" } ], "output_type": "TEXT" } ] } ] } ``` Each bridge block represents one input class field. In this example, the first bridge exposes **Loan\_Amount** from **Loan\_Application** and the second exposes **Wages\_Tips** from **W2**. For each bridge, three values are project-specific and must be updated together: `label` (the stable identifier downstream formulas reference), `class_name`, and `field_name`. Add one bridge block for each `(class, field)` pair your cross-class fields depend on. Append cross-class field columns after all the bridges. The `class_fields` value on `get_input_fields_per_class_field` is project-specific -- replace `{}` with your own class-to-field mapping (such as `{"Loan_Application": ["Loan__Amount"], "W2": ["Wages__Tips"]}`). The `kwargs` input on each bridge line is required by the refiner schema but ignored at runtime; keep `"value": "{}"`. --- **Pattern B: `get_all_intra_class_fields` + `get_intra_class_field_from_calculated` (two-step)** An alternative pattern uses a single shared aggregation column (`get_all_intra_class_fields`) followed by individual bridge columns that each call `get_intra_class_field_from_calculated` to extract one `(class, field)` pair from it. The `field_selection_strategy` is set one time on the aggregation column rather than per bridge. ```json { "dev_input": {}, "options": { "provenance_tracking": true, "auto_provenance": false }, "fields": [ { "label": "__class_field_inputs", "lines": [ { "function_id": { "name": "get_input_fields_per_class_field", "source": "NATIVE" }, "inputs": [ { "arg_name": "class_fields", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "ANY", "data_type_options": ["TEXT"], "value": "{}" } ] } ] }, { "label": "__all_intra_class_fields", "lines": [ { "function_id": { "name": "get_all_intra_class_fields", "source": "NATIVE" }, "inputs": [ { "arg_name": "class_field_inputs", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "FIELD", "data_type_options": ["TEXT"], "value": "__class_field_inputs" }, { "arg_name": "field_selection_strategy", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "ANY", "data_type_options": ["TEXT"], "value": "\"best_field_value\"" }, { "arg_name": "kwargs", "arg_type": "VAR_KEYWORD", "data_type": "DICT", "data_type_options": ["DICT"], "value": "{}" } ], "output_type": "TEXT" } ] }, { "label": "__Loan__Application_Loan__Amount_", "lines": [ { "function_id": { "name": "get_intra_class_field_from_calculated", "source": "NATIVE" }, "inputs": [ { "arg_name": "all_intra_class_fields", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "FIELD", "data_type_options": ["TEXT"], "value": "__all_intra_class_fields" }, { "arg_name": "class_name", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "ANY", "data_type_options": ["TEXT"], "value": "\"Loan_Application\"" }, { "arg_name": "field_name", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "ANY", "data_type_options": ["TEXT"], "value": "\"Loan__Amount\"" }, { "arg_name": "kwargs", "arg_type": "VAR_KEYWORD", "data_type": "DICT", "data_type_options": ["DICT"], "value": "{}" } ], "output_type": "TEXT" } ] }, { "label": "__W2_Wages__Tips_", "lines": [ { "function_id": { "name": "get_intra_class_field_from_calculated", "source": "NATIVE" }, "inputs": [ { "arg_name": "all_intra_class_fields", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "FIELD", "data_type_options": ["TEXT"], "value": "__all_intra_class_fields" }, { "arg_name": "class_name", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "ANY", "data_type_options": ["TEXT"], "value": "\"W2\"" }, { "arg_name": "field_name", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "ANY", "data_type_options": ["TEXT"], "value": "\"Wages__Tips\"" }, { "arg_name": "kwargs", "arg_type": "VAR_KEYWORD", "data_type": "DICT", "data_type_options": ["DICT"], "value": "{}" } ], "output_type": "TEXT" } ] } ] } ``` In this pattern, `kwargs` on `get_all_intra_class_fields` and each `get_intra_class_field_from_calculated` line is required by the refiner schema but ignored at runtime; keep `"value": "{}"`. For `custom_udf` strategies, pass an extra `udf_json` argument on `get_all_intra_class_fields`. #### Derived cross-class fields (`choice_type`: DERIVED) Prompt-based columns use `refine_field_v2`. The first prompt line sets `previous_line` to `null`; later lines chain `previous_line` to the prior line identifier (such as `Is_Reasonable_Loan@0`). `prompt_arg_vals` reference bridge label strings. #### Derived field example (Is\_Reasonable\_Loan) Three inputs use project-specific values that must be updated together: the prompt string in `prompt` (which embeds human-readable class field labels), `prompt_arg_names` (which lists those same human-readable labels), and `prompt_arg_vals` (which references the corresponding bridge labels from the fixed prefix). In this example, the field asks whether a loan amount from **Loan\_Application** is reasonable given the wages from **W2**. ```json { "label": "Is_Reasonable_Loan", "lines": [ { "function_id": { "name": "refine_field_v2", "source": "NATIVE" }, "inputs": [ { "arg_name": "prompt", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "TEXT", "data_type_options": ["TEXT"], "value": "\"Is giving a loan of \\\\ Loan_Application.Loan_Amount \\\\ reasonable to someone making \\\\ W2.Wages_Tips \\\\ before taxes per year? Return in 1 word, Yes or No\"" }, { "arg_name": "metrics_params", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "ANY", "data_type_options": ["TEXT"], "value": "{\"source_id\": \"\", \"source_type\": \"\", \"user_agent\": \"\"}" }, { "arg_name": "proj_uuid", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "TEXT", "data_type_options": ["TEXT"], "value": "\"\"" }, { "arg_name": "cached_index", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "TEXT", "data_type_options": ["TEXT"], "value": "null", "default": "null" }, { "arg_name": "custom_request", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "ANY", "data_type_options": ["TEXT"], "value": "{}" }, { "arg_name": "doc_id", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "TEXT", "data_type_options": ["TEXT"], "value": "null", "default": "null" }, { "arg_name": "data_type", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "TEXT", "data_type_options": ["TEXT"], "value": "\"TEXT\"" }, { "arg_name": "previous_line", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "TEXT", "data_type_options": ["TEXT"], "value": "null", "default": "null" }, { "arg_name": "prompt_arg_names", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "LIST", "data_type_options": ["TEXT"], "value": "[{\"data_type\": \"ANY\", \"value\": \"\\\"Loan_Application.Loan_Amount\\\"\"}, {\"data_type\": \"ANY\", \"value\": \"\\\"W2.Wages_Tips\\\"\"}]" }, { "arg_name": "prompt_arg_vals", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "LIST", "data_type_options": ["TEXT"], "value": "[{\"data_type\": \"FIELD\", \"value\": \"__Loan__Application_Loan__Amount_\"}, {\"data_type\": \"FIELD\", \"value\": \"__W2_Wages__Tips_\"}]" }, { "arg_name": "prompt_arg_data_types", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "LIST", "data_type_options": ["TEXT"], "value": "[{\"data_type\": \"ANY\", \"value\": \"\\\"TEXT\\\"\"}, {\"data_type\": \"ANY\", \"value\": \"\\\"TEXT\\\"\"}]" }, { "arg_name": "kwargs", "arg_type": "VAR_KEYWORD", "data_type": "DICT", "data_type_options": ["DICT"], "value": "{}" } ], "output_type": "TEXT" } ] } ``` #### Ranked/aggregated cross-class fields (`choice_type`: AGGREGATED) `selection_logic` (such as `FIRST_VALID`) maps to the first line's native function name in lowercase (`first_valid`). `input_fields` order drives `input_field_names` and `input_field_vals`. The second line is always `ranked_value_unwrap`, referencing the first line by label and index (such as `Applicant_Name@0`). Other ranking natives follow the same two-line pattern with a different `function_id.name`: `ocr_confidence`, `field_confidence`, `first_ranked`. #### Aggregated field example (Applicant\_Name) This example picks the applicant name by trying each class in priority order -- **Loan\_Application**, then **Pay\_Stub**, then **Bank\_Statement**, then **W2** -- and returning the first valid value. `input_field_names` and `input_field_vals` must be listed in the same order and reference bridge labels from the fixed prefix. ```json { "label": "Applicant_Name", "lines": [ { "function_id": { "name": "first_valid", "source": "NATIVE" }, "inputs": [ { "arg_name": "input_field_names", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "LIST", "data_type_options": ["TEXT"], "value": "[{\"data_type\": \"ANY\", \"value\": \"\\\"__Loan__Application_Borrower__Name_\\\"\"}, {\"data_type\": \"ANY\", \"value\": \"\\\"__Pay__Stub_Employee_Name_\\\"\"}, {\"data_type\": \"ANY\", \"value\": \"\\\"__Bank__Statement_Account__Holder_\\\"\"}, {\"data_type\": \"ANY\", \"value\": \"\\\"__W2_Employer__Name_\\\"\"}]" }, { "arg_name": "input_field_vals", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "LIST", "data_type_options": ["TEXT"], "value": "[{\"data_type\": \"FIELD\", \"value\": \"__Loan__Application_Borrower__Name_\"}, {\"data_type\": \"FIELD\", \"value\": \"__Pay__Stub_Employee_Name_\"}, {\"data_type\": \"FIELD\", \"value\": \"__Bank__Statement_Account__Holder_\"}, {\"data_type\": \"FIELD\", \"value\": \"__W2_Employer__Name_\"}]" }, { "arg_name": "kwargs", "arg_type": "VAR_KEYWORD", "data_type": "DICT", "data_type_options": ["DICT"], "value": "{}" } ], "output_type": "TEXT" }, { "function_id": { "name": "ranked_value_unwrap", "source": "NATIVE" }, "inputs": [ { "arg_name": "ranked_value", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "LINE", "data_type_options": ["TEXT"], "value": "Applicant_Name@0" }, { "arg_name": "kwargs", "arg_type": "VAR_KEYWORD", "data_type": "DICT", "data_type_options": ["DICT"], "value": "{}" } ] } ] } ``` #### Custom function cross-class fields (`choice_type`: UDF) The primary line calls `run_case_udf_in_lambda_v2` with: script path, serialized `function_args`, `udf_id`, `lambda_udf_id`, generated `code_arg` and `fn_name`, a `class_fields` map of human-readable names to bridge labels, `udf_imports`, and `args` for the lambda. ENUM fields may add `select_options` and `output_type: "SELECT"` on the column object. #### UDF field example (user\_\_has\_\_enough\_\_money\_\_in\_\_bank) This example calls a Python function that checks whether the loan amount is less than ten times the bank account's ending balance. It takes `Is_Reasonable_Loan` -- a cross-class field computed earlier in the same program -- as a positional input via `args`. `class_fields` maps human-readable `ClassName.FieldName` keys to their corresponding bridge labels, making those values available to the function via the `class_fields` parameter. `function_args` lists any cross-class fields passed as named inputs. ```json { "label": "user__has__enough__money__in__bank", "lines": [ { "function_id": { "name": "run_case_udf_in_lambda_v2", "source": "NATIVE" }, "inputs": [ { "arg_name": "absolute_scripts_path", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "TEXT", "data_type_options": ["TEXT"], "value": "\"\"" }, { "arg_name": "function_args", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "TEXT", "data_type_options": ["TEXT"], "value": "[{\"name\": \"Is_Reasonable_Loan\", \"data_type\": \"FIELD\", \"value\": \"Is_Reasonable_Loan\"}]" }, { "arg_name": "lambda_udf_id", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "TEXT", "data_type_options": ["TEXT"], "value": "\"\"" }, { "arg_name": "packet_id", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "TEXT", "data_type_options": ["TEXT"], "value": "null", "default": "null" }, { "arg_name": "udf_id", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "TEXT", "data_type_options": ["TEXT"], "value": "\"\"" }, { "arg_name": "code_arg", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "TEXT", "data_type_options": ["TEXT"], "value": "\"\"" }, { "arg_name": "fn_name", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "TEXT", "data_type_options": ["TEXT"], "value": "\"user_has_enough_money_in_bank\"" }, { "arg_name": "class_fields", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "DICT", "data_type_options": ["TEXT"], "value": "{\"Loan_Application.Loan_Amount\": {\"name\": \"Loan_Application.Loan_Amount\", \"data_type\": \"FIELD\", \"value\": \"__Loan__Application_Loan__Amount_\"}, \"Bank_Statement.Ending_Balance\": {\"name\": \"Bank_Statement.Ending_Balance\", \"data_type\": \"FIELD\", \"value\": \"__Bank__Statement_Ending__Balance_\"}}" }, { "arg_name": "udf_imports", "arg_type": "POSITIONAL_OR_KEYWORD", "data_type": "TEXT", "data_type_options": ["TEXT"], "value": "[]" }, { "arg_name": "args", "arg_type": "VAR_POSITIONAL", "data_type": "LIST", "value": "[{\"data_type\": \"FIELD\", \"value\": \"Is_Reasonable_Loan\"}]" }, { "arg_name": "kwargs", "arg_type": "VAR_KEYWORD", "data_type": "DICT", "data_type_options": ["DICT"], "value": "{}" } ], "output_type": "TEXT" } ] } ``` **Column ordering** -- Emit fixed-prefix columns first, then cross-class field columns in project order. ## Cross-class checkpoint This step is optional and applies only when the published project defines validation rules scoped to cross-class fields (`case_field_validation_rules` non-empty at publish). It runs after *Process case* on the case step's output folder, so rules see both class extraction and materialized cross-class fields. The checkpoint uses a [validations module](/flow/step-config-reference/creating-validation-checkpoints). ``` modules/case-field-validations.checkpoint/validations.ibvalidations ``` This validation file embeds cross-class settings from the project (`project.settings.cross_class`), so rules behave consistently with the app editor experience. Supported rule types include `ALL_INPUTS_VALID`, `ALL_INPUTS_MATCH`, `CHOSEN_FIELD_VALID`, and `UDF`. | Kwarg | Purpose | | ------------------------- | --------------------------------------------------------------------- | | `input_folder` | Output folder of the *Process case* step. | | `output_folder` | Checkpoint output for downstream steps. | | `ibvalidations_path` | `modules/case-field-validations.checkpoint/validations.ibvalidations` | | `validations_module` | Same path (duplicated for flow UI or executor). | | `settings.enable_stp` | Typically `"false"`. | | `settings.execution_mode` | `"run_validations_only"` | ```json { "kwargs": { "input_folder": "s5_case", "output_folder": "s6_apply_checkpoint", "ibvalidations_path": "modules/case-field-validations.checkpoint/validations.ibvalidations", "validations_module": "modules/case-field-validations.checkpoint/validations.ibvalidations", "settings": { "enable_stp": "false", "execution_mode": "run_validations_only" } }, "option": { "internalName": "apply_checkpoint", "name": "Apply Checkpoint" } } ``` Validation file top-level shape: ```json { "mode": "case", "metadata": {}, "validation_rules": [] } ``` This module is distinct from class-level checkpoints such as `modules/all-classes.checkpoint/validations.ibvalidations`, which run earlier on single-class extraction. ## Combine all Use this final *Combine* step when you need all classes and packet-level fields together in one stream for downstream steps. | Setting | Value | | --------------------------- | ----- | | `combine_same_file_records` | `no` | | `combine_all_file_records` | `no` | > Design flows for packet processing.