Skip to main content

Set Up, Scale, and Run an Extraction

Learning Objectives

After completing this unit, you’ll be able to:

  • Describe the setup an admin completes before extraction can run.
  • Explain how batch processing enables high volume extraction.
  • Explain what happens when a document is extracted.
  • Recognize how a few new records track each run on top of the platform's existing context and AI services.

Configure and Trace an Extraction

In the previous unit, you saw the user workflow for processing a patient referral. That workflow depends on configuration completed before the document is uploaded and on a set of platform records and services that manage each extraction.

An admin first selects or configures a template for the document type. The template defines the entities and attributes to extract, the prompts that guide identification, the mappings to Salesforce records and fields, and the confidence threshold used during review.

When a document is processed, Salesforce creates a Document Extraction Request and tracks the extraction through its component steps. Those steps capture the source text, structure the extracted values, resolve references and record matches, and prepare the confirmed results for saving.

In this unit, you learn how admins configure extraction templates and batch processing, and how Salesforce records and executes each stage of a document extraction.

Start with a Standard Template

Each document type requires a defined extraction structure. An admin must specify which entities and attributes to identify, how those values map to Salesforce records and fields, and which prompts and settings guide the extraction.

Document AI for Health provides standard templates for common healthcare documents so admins don’t have to build that configuration from scratch. Each template is paired with a context definition and context mapping that together describe what to extract and where the reviewed values belong.

On the Document AI for Health setup page, admins can start with templates for patient referrals, disease definitions, and case intake forms.

Document AI for Health setup with Extraction Templates.

At Cumulus Health, the Patient Referral template supports the workflow Dana used in the previous unit. Its preconfigured structure gives the extraction process a consistent set of entities, attributes, and mappings for that document type.

Admins can use a standard template when it fits the organization’s process. For document types outside the standard set, they can create a custom template and define the required extraction structure themselves.

Configure How a Template Extracts Data

The standard Patient Referral template brings together the configuration required to process that document type. It references a prebuilt context definition and context mapping, selects the large language model used for extraction, and provides the instructions and review settings that govern the run.

Start with the context definition, which models the information the extraction should return. It organizes that structure into nodes, attributes, and relationships.

In the Patient Referral context definition, nodes represent entities such as the patient account, the clinical service request that anchors the referral, and related clinical information. Attributes describe the values associated with those nodes. The account node, for example, includes attributes for the patient’s first name, birthdate, and gender.

The Patient Referral context definition structure including the nodes for each object, with the attributes listed under the Account node.

The associated context mapping connects that modeled structure to the Health Cloud data model. It links nodes and attributes from the context definition to the Salesforce objects and fields that receive the reviewed extraction results. The patient birthdate attribute, for example, maps to the corresponding field on the account record.

The context definition and mapping establish the structure and destination of the extracted data. The extraction template then configures how the feature identifies that data in a particular document type.

An extraction template includes:

  • The large language model used to process the document.
  • A description that provides document-level instructions and context for the extraction.
  • Field prompts that tell the model what each configured attribute represents and what information to identify.
  • A low-confidence threshold that determines which extracted values are flagged for additional review.
  • Optional post-transform and post-save flows that extend processing at defined points in the run.

On the template page, you can view the associated context definitions and mappings, extraction flows, and field mapping prompts.

The Patient Referral template's configuration.

The field prompts are especially important because the context definition names the expected attributes but doesn’t explain how to recognize their values in the source document. In the Patient Referral template, the prompt for a health condition instructs the model to identify each distinct condition named in the referral, including conditions presented in a list or within narrative text.

The template-level description provides broader guidance for interpreting the document as a whole. Field prompts narrow that guidance for individual attributes. Together, they give the selected LLM the instructions to propose values for the structure supplied by the context definition.

The template also sets the low-confidence threshold. The Patient Referral template uses a threshold of 80, so any extracted field scored below 80 is flagged for additional manual review. The score directs reviewer attention; it doesn’t establish that higher-scoring values are clinically correct.

The Patient Referral template's details, including the model, the plain-language description, and confidence threshold.

With the standard Patient Referral template, Salesforce also provides the context definition and the mapping it references. For another document type, an admin can create the required context structure and mapping, associate them with a custom extraction template, write or refine the prompts, test the results, and activate the template.

Set Up Batch Processing

The Patient Referral template can support both individually submitted documents and recurring, high-volume intake. The template still defines how each referral is extracted; batch processing changes how files enter that extraction workflow.

In Batch Document AI setup, an admin creates a batch configuration and:

  • Selects the extraction template to apply.
  • Identifies the connected folder that supplies the documents.
  • Assigns an owner for the resulting extraction requests.
  • Sets the schedule for checking and processing new files.
  • Chooses whether each extraction pauses for review before saving.

Batch Document AI setup, with a list of templates, along with their respective statuses and folder locations.

A batch job can run on a recurring schedule as frequently as every 15 minutes. When a supported document is added to the configured folder, the job submits it for processing and creates a separate document extraction request for that file.

Each document then moves through the same extraction stages used for an individually submitted request. Depending on the batch configuration, the result can pause for a user to review extracted values and resolve record matches before Salesforce saves the records.

This way, batch processing automates the repeated intake step. Teams don’t need to open the extraction interface and initiate a separate request for every incoming document. Meanwhile, each file is still individually tracked for review, troubleshooting, and audit.

Trace a Run from Document to Records

Batch processing changes how documents enter the workflow. Whether a user starts an extraction directly or a scheduled job submits it from a connected folder, Salesforce creates a Document Extraction Request for that document.

The request retains the source file and records the status of the run from extraction through saving. The related Document Extraction Request Step records track the processing stages completed for that request.

A Document Extraction Request with its child steps, Extract, Transform, and Post-Transform Action, all completed.

The batch run includes three core stages.

  • Extract: Optical character recognition captures the source text, and the selected model uses the template instructions to propose values for the configured attributes. Each extracted value includes a confidence score.
  • Transform: Salesforce organizes those proposed values according to the referenced context definition and context mapping, producing the target record structure for the document type.
  • Post-transform action: Salesforce resolves references and applies matching logic so extracted entities can connect to records already in the system instead of automatically creating duplicates.

After those stages are completed, the request can pause for review. A user confirms or corrects the extracted values, resolves record matches, and saves the approved results. Salesforce then creates or updates the connected records, and the request status changes to Save Completed.

Because the request and its steps preserve the status of each stage, admins can inspect failed runs, review processing history, and rerun an extraction when needed.

Wrap Up

Document AI for Health separates document-processing configuration from day-to-day intake. Admins define how each document type is processed through templates, associated context resources, prompts, mappings, thresholds, and batch settings. Intake teams then use that configuration to extract, review, match, and save document information without rebuilding the process for each file.

Every document is individually tracked through a Document Extraction Request and its processing steps, whether the run begins with a direct upload or a scheduled batch job. That gives organizations like Cumulus a consistent way to process common health documents while preserving human review, record matching, and execution history.

To learn more about the underlying Document AI services in Data 360 and how generative AI requests pass through the Einstein Trust Layer, see the Resources.

Resources

Salesforce 도움말에서 Trailhead 피드백을 공유하세요.

Trailhead에 관한 여러분의 의견에 귀 기울이겠습니다. 이제 Salesforce 도움말 사이트에서 언제든지 새로운 피드백 양식을 작성할 수 있습니다.

자세히 알아보기 의견 공유하기