Skip to main content

Get Started with Document AI

Learning Objectives

After completing this unit, you’ll be able to:

  • Explain what Document AI for Health is and how it extracts data from unstructured health documents.
  • Distinguish generative AI extraction from traditional optical character recognition.
  • Describe the core capabilities and templates of Document AI for Health.
  • Explain how Document AI for Health fits within Health Cloud and Data 360.

Before You Start

Before you start this module, consider completing this recommended content.

Address the Document Burden in Health

Health organizations receive critical information through documents every day. Patient referrals, case intake forms, disease definitions, lab results, prior authorization materials, and clinical correspondence can arrive as emailed PDFs, portal uploads, scanned packets, or faxes.

Even information inside digital files often isn’t structured for use in a health system. A referral coordinator can still need to identify the patient, confirm the referring provider, interpret the requested service, review coverage information, and enter each detail into the appropriate records. Similar work occurs when teams process intake forms or prepare disease information for surveillance workflows.

That manual intake creates more than an administrative burden. A referral can’t move toward review and scheduling until its information is captured correctly. A case can’t enter the appropriate workflow until staff classify and record its details. And information contained in a document can’t reliably drive automation, reporting, or coordinated follow-up while it remains only text in an attached file.

Document AI for Health helps close that gap. It extracts information from unstructured health documents, organizes the results according to a defined data structure, and maps them to Salesforce records for review and use.

In this badge, you learn how Document AI for Health processes a document from upload through extraction, review, matching, and saving. You also explore the templates, context definitions, mappings, and batch-processing options that admins use to support different document types and intake volumes.

Meet Document AI for Health

Document AI for Health turns information contained in health documents into structured Salesforce data. It uses optical character recognition (OCR) to capture text from PDFs and images, and then uses generative AI to identify the values defined by an extraction template.

The template specifies what information to extract and how to map it to Health Cloud objects and fields. For a patient referral, for example, the feature can identify patient details, the referring provider, the clinical reason, coverage information, and the requested service. It returns those values with confidence scores so a user can review the results, match them to existing records, and save the confirmed data.

Document AI for Health includes standard templates for common healthcare documents, mappings to the Health Cloud data model, and records that track each extraction. Users can easily review the extraction results to resolve low-confidence values and record mismatches.

This approach differs from traditional OCR, which converts scanned or photographed content into machine-readable text, but can’t determine which text represents a patient name, referral date, diagnosis, or referring provider.

Document AI for Health uses the text, its surrounding context, and the instructions in the selected template to identify those values without relying on one fixed page layout. A referring provider might appear beneath a labeled field in one document, in a letterhead block in another, and within a paragraph in a third. The template defines what the model should identify, while its mappings define where the reviewed values belong.

Use Document Data Across Health Cloud

Because Document AI for Health writes to the shared Health Cloud data model, organizations can use extracted information across the broader Health Cloud portfolio.

A patient referral can supply records used in care coordination. A case intake form can populate information required for a case-management workflow. A disease-definition document can provide structured criteria for disease surveillance. The document type and selected template change, but the underlying pattern stays consistent: unstructured content becomes Health Cloud data that the relevant application can use.

This architecture diagram shows the Health Cloud ecosystem that Document AI for Health operates in.Health Cloud with Data 360 at the foundation and Health Cloud applications above it.

Data 360 provides the underlying data and AI services, while Health Cloud applications use the resulting data across healthcare, home health, care coordination, appointment management, health operations, disease surveillance, and other solution areas. Agentforce agents and Slack can also work with that shared data in the context where teams already operate.

Explore the Core Capabilities

Document AI for Health combines document extraction with health-specific configuration, review, and record handling. Its core capabilities include:

  • Standard and custom extraction templates: Use prebuilt templates for patient referrals, case intake, and disease definitions, or configure a template for another document type. Each template defines the information to extract and how the results map to Salesforce records and fields.
  • Field-level confidence scoring: Review a confidence score for each extracted value and use a configured threshold to identify results that need closer attention before saving documents.
  • Support for common document formats: Process PDFs and image files such as JPEG and PNG, including scans, photographs, and supported handwritten content.
  • Record matching and reference resolution: Match extracted information to existing records, resolve related references, and avoid creating unnecessary duplicates when you review and save documents.
  • Batch extraction: Run extraction against documents placed in a connected folder on a recurring schedule. Admins can also require a review before records are saved.
  • Amazon S3 document source: Read files for extraction directly from a connected Amazon S3 bucket, for both single-document and batch runs.
  • Automatic template matching: Identify the right template for an incoming document, without manual selection.

Together, these capabilities reduce the manual work between document receipt and record creation. Staff still review low-confidence values and confirm record matches, but they no longer need to locate and reenter every detail from the source document.

Follow the Path from Upload to Record

Before a team can process a document, an admin selects or configures an extraction template for that document type. The template defines what information to identify, how to structure it, and where the reviewed values map in Salesforce. You explore that setup in the final unit.

Once the configuration is in place, each document follows the same general path.

Stage

What Happens

Upload

A user uploads a supported PDF or image, or a batch job submits a file from a connected folder.

Extract

Optical character recognition captures the document text. The selected model then identifies the entities and values defined by the extraction template and assigns a confidence score to each result.

Transform

The platform organizes the extracted values into the record structure defined by the template and its mappings.

Review and Match

A user compares the proposed values with the source document, corrects low-confidence results, and matches the extracted information to existing records or confirms that new records should be created.

Save

Salesforce creates or updates the connected records and stores the completed extraction request.

The platform can also run configurable steps after the Transform or Save stages, where an organization can insert its own logic, such as enriching a record once it's created. You learn about these later in this badge.

By the end of the process, the document’s contents are available as structured records that can enter the relevant health workflow. Staff review the proposed results rather than transcribing the source document field by field.

What’s Next?

Document AI for Health turns information from health documents into structured records while keeping review and record matching in the workflow for data consistency and accuracy.

At Cumulus Health, referral coordinator Dana Okafor (she/her) receives patient referrals from external providers throughout the day. Each referral contains information that must be reviewed, matched to the correct patient and related records, and made available to the care team.

In the next unit, you follow Dana as she uploads a referral, reviews the extracted values and confidence scores, resolves record matches, and saves the completed referral to Salesforce.

Resources

Salesforce ヘルプで Trailhead のフィードバックを共有してください。

Trailhead についての感想をお聞かせください。[Salesforce ヘルプ] サイトから新しいフィードバックフォームにいつでもアクセスできるようになりました。

詳細はこちら フィードバックの共有に進む