[!NOTE] TL;DR: Build a serverless Document AI pipeline on AWS using S3, Textract, Bedrock Converse, and DynamoDB orchestrated by Step Functions. Extract typed records with Claude 3.5 Haiku while slashing extraction costs by 80% compared to native Textract forms. Complete deployable Python code with Pydantic validation included below.
The broken invoice batch
A thousand supplier invoices flooded my S3 bucket last Tuesday. My Lambda promptly choked, throwing memory errors with the theatrical flair of an opera diva. I spent three hours apologizing to our accounting director before inspecting the logs.
Every single PDF exceeded twenty megabytes. The synchronous endpoint refused the workload.
Building on this incident, I re-architected our intake from scratch. This tutorial details the resulting build: a production-grade Document AI pipeline on AWS that turns raw scanned files into typed records without burning budget.
The architectural bottleneck
Engineers instinctively reach for Textract built-in forms and tables features when parsing business documents. One spontaneously imagines that proprietary AWS layout detection solves structural extraction cleanly. The numbers say the opposite.
According to the official AWS Textract pricing, raw text detection (DetectDocumentText and asynchronous StartDocumentTextDetection) costs 1.50 USD per 1,000 pages. The moment you activate key-value forms and tables via AnalyzeDocument, the bill leaps to 65.00 USD per 1,000 pages. That is a 43-fold price increase. For an organisation processing 50,000 pages monthly, the extraction bill jumps from 75 USD to 3,250 USD.
The cost of native Textract key-value extraction (and this is where most architectures burn cash needlessly) delivers rigid, brittle key pairs that break whenever a supplier shifts a table border by two millimetres.
Think of raw OCR as surveying the physical footprint of an old building. The surveyor maps every brick, doorway, and lintel without judging the tenant furniture. If you ask the surveyor to catalogue every antique vase on site, costs skyrocket and errors multiply. Let the surveyor map the perimeter stone, then send an interior architect to interpret the rooms.
My main question was how to combine low-cost OCR with large language model reasoning inside an asynchronous, event-driven state machine.
flowchart LR
S3[S3 Ingestion Bucket] --> EB[Amazon EventBridge]
EB --> SF[Step Functions DPIPE]
SF --> L1[Lambda: Start Textract]
L1 --> TX[Amazon Textract]
TX --> SF
SF --> L2[Lambda: Bedrock Extract]
L2 --> BD[Bedrock Converse API]
BD --> L2
L2 --> DDB[(Amazon DynamoDB)]
SF --> CW[CloudWatch Logs and Metrics]
Target architecture and specifications
Here I present the DPIPE blueprint. The pipeline processes multi-page documents asynchronously without holding expensive Lambda execution hours.
Four functional pillars form DPIPE: 1. Ingestion: an S3 bucket emits an EventBridge event upon object creation. 2. Character extraction: an asynchronous Textract job runs raw OCR at 1.50 USD per 1,000 pages. 3. Semantic normalization: Claude 3.5 Haiku on Amazon Bedrock receives linearized text and extracts structured data against a strict Pydantic schema. 4. Persistence and audit: DynamoDB stores typed records with automated TTL and validation statuses.
AWS Step Functions coordinates the state transitions (without keeping Lambda execution timers running during OCR), saving substantial compute charges. We use AWS Step Functions standard workflows, billed at 0.025 USD per 1,000 state transitions.
Step 1: Storage and event notification
We begin by establishing an isolated S3 ingestion bucket with EventBridge notifications enabled.
Run this AWS CLI command to configure your bucket:
aws s3api create-bucket --bucket dpipe-invoices-ingest-us-east-1 --region us-east-1
aws s3api put-bucket-notification-configuration --bucket dpipe-invoices-ingest-us-east-1 --notification-configuration '{"EventBridgeConfiguration": {}}'
Expected result: S3 publishes an event to the default EventBridge bus for every PutObject call, decoupling uploads from pipeline execution.
Step 2: Asynchronous text detection with Textract
Synchronous Textract calls fail on multi-page PDFs over 10 MB. Asynchronous jobs via StartDocumentTextDetection support files up to 500 MB and 3,000 pages.
The first Lambda function initiates text detection and returns the job identifier to the state machine:
import os
import boto3
from pydantic import BaseModel, Field
textract_client = boto3.client("textract", region_name=os.environ.get("AWS_REGION", "us-east-1"))
class StartOcrRequest(BaseModel):
model_config = {"strict": True}
bucket_name: str = Field(..., description="S3 input bucket name.")
object_key: str = Field(..., description="S3 input object key.")
class StartOcrResponse(BaseModel):
model_config = {"strict": True}
job_id: str = Field(..., description="Textract asynchronous job identifier.")
status: str = Field(default="IN_PROGRESS", description="Initial job status.")
def lambda_handler(event: dict, context: object) -> dict:
payload = StartOcrRequest.model_validate(event)
response = textract_client.start_document_text_detection(
DocumentLocation={"S3Object": {"Bucket": payload.bucket_name, "Name": payload.object_key}}
)
result = StartOcrResponse(job_id=response["JobId"])
return result.model_dump()
Expected result: Textract registers the job and returns a unique JobId within 400 milliseconds.
Step 3: Structured field normalization with Bedrock
Once Textract finishes, we fetch the detected text lines and pass them to Amazon Bedrock.
We use the Amazon Bedrock Converse API with anthropic.claude-3-5-haiku-20241022-v1:0. At 0.80 USD per 1M input tokens and 4.00 USD per 1M output tokens, Haiku costs a fraction of Sonnet while delivering high precision on structured extraction.
We define our invoice schema with Pydantic, enforcing type coercion and field validation:
import json
import boto3
from typing import List, Optional
from pydantic import BaseModel, Field, field_validator
bedrock_client = boto3.client("bedrock-runtime", region_name="us-east-1")
class LineItem(BaseModel):
model_config = {"strict": True}
description: str = Field(..., description="Product or service description.")
quantity: float = Field(..., gt=0, description="Item quantity.")
unit_price: float = Field(..., ge=0, description="Unit price before tax.")
total_price: float = Field(..., ge=0, description="Calculated line total.")
class InvoiceDocument(BaseModel):
model_config = {"strict": True}
invoice_number: str = Field(..., description="Unique invoice or reference code.")
vendor_name: str = Field(..., description="Issuing supplier company name.")
invoice_date: str = Field(..., description="Invoice issue date in YYYY-MM-DD format.")
currency: str = Field(default="USD", description="Three-letter currency code.")
line_items: List[LineItem] = Field(..., description="List of invoiced items.")
total_amount: float = Field(..., gt=0, description="Grand total including taxes.")
@field_validator("currency")
@classmethod
def validate_currency(cls, value: str) -> str:
if len(value) != 3:
raise ValueError("Currency must be a 3-letter ISO code.")
return value.upper()
def extract_invoice_fields(raw_text: str) -> InvoiceDocument:
tool_spec = {
"tools": [{
"toolSpec": {
"name": "record_invoice",
"description": "Persist extracted structured invoice information.",
"inputSchema": {"json": InvoiceDocument.model_json_schema()}
}
}],
"toolChoice": {"tool": {"name": "record_invoice"}}
}
messages = [{"role": "user", "content": [{"text": f"Extract structured invoice data from this OCR text:\n\n{raw_text}"}]}]
response = bedrock_client.converse(
modelId="anthropic.claude-3-5-haiku-20241022-v1:0",
messages=messages,
toolConfig=tool_spec,
inferenceConfig={"temperature": 0.0, "maxTokens": 2048}
)
content = response["output"]["message"]["content"]
tool_use = next(block["toolUse"] for block in content if "toolUse" in block)
return InvoiceDocument.model_validate(tool_use["input"])
Expected result: Claude 3.5 Haiku returns arguments constrained strictly to the JSON schema, and Pydantic validates the parsed dictionary directly into an InvoiceDocument instance.
Step 4: DynamoDB persistence and schema validation
The validated record must be persisted with execution metadata and audit timestamps.
Here is the persistence logic writing directly to Amazon DynamoDB:
import os
import boto3
from datetime import datetime, timezone
dynamodb = boto3.resource("dynamodb", region_name=os.environ.get("AWS_REGION", "us-east-1"))
table = dynamodb.Table(os.environ.get("TABLE_NAME", "dpipe_extracted_documents"))
def save_document(invoice: InvoiceDocument, s3_uri: str, job_id: str) -> dict:
item = {
"pk": f"VENDOR#{invoice.vendor_name.upper().replace(' ', '_')}",
"sk": f"INVOICE#{invoice.invoice_number}",
"vendor_name": invoice.vendor_name,
"invoice_date": invoice.invoice_date,
"currency": invoice.currency,
"total_amount": str(invoice.total_amount),
"line_items": [item.model_dump() for item in invoice.line_items],
"source_s3_uri": s3_uri,
"textract_job_id": job_id,
"created_at": datetime.now(timezone.utc).isoformat(),
"status": "PROCESSED"
}
table.put_item(Item=item)
return {"status": "SUCCESS", "document_id": f"{item['pk']}#{item['sk']}"}
Expected result: DynamoDB stores the item under a partitioned vendor key for rapid querying and accounting audits.
Step 5: Step Functions state machine assembly
Step Functions ties these components into a resilient state machine with exponential retries and automated polling.
Here is the Amazon States Language definition:
Comment: "DPIPE Serverless Document AI Pipeline"
StartAt: StartTextractOcr
States:
StartTextractOcr:
Type: Task
Resource: "arn:aws:lambda:us-east-1:123456789012:function:dpipe-start-ocr"
Next: WaitForTextract
WaitForTextract:
Type: Wait
Seconds: 15
Next: CheckTextractStatus
CheckTextractStatus:
Type: Task
Resource: "arn:aws:lambda:us-east-1:123456789012:function:dpipe-check-ocr"
Next: EvaluateOcrCondition
EvaluateOcrCondition:
Type: Choice
Choices:
- Variable: "$.status"
StringEquals: "SUCCEEDED"
Next: ExtractWithBedrock
- Variable: "$.status"
StringEquals: "IN_PROGRESS"
Next: WaitForTextract
Default: PipelineFailed
ExtractWithBedrock:
Type: Task
Resource: "arn:aws:lambda:us-east-1:123456789012:function:dpipe-bedrock-extract"
Retry:
- ErrorEquals: ["BedrockThrottlingException", "States.TaskFailed"]
IntervalSeconds: 3
MaxAttempts: 3
BackoffRate: 2.0
Next: PipelineCompleted
PipelineCompleted:
Type: Succeed
PipelineFailed:
Type: Fail
Cause: "Textract processing failed or was cancelled."
Expected result: Step Functions polls Textract every 15 seconds without consuming Lambda runtime while the OCR engine reads document pages.
The complete runnable pipeline
To test the entire workflow locally or inside an AWS Lambda package, here is the integrated handler combining pagination retrieval, text reconstruction, Bedrock inference, and storage:
import os
import boto3
from typing import List
from pydantic import BaseModel, Field, field_validator
textract_client = boto3.client("textract", region_name="us-east-1")
bedrock_client = boto3.client("bedrock-runtime", region_name="us-east-1")
dynamodb = boto3.resource("dynamodb", region_name="us-east-1")
table = dynamodb.Table(os.environ.get("TABLE_NAME", "dpipe_extracted_documents"))
class LineItem(BaseModel):
model_config = {"strict": True}
description: str = Field(..., description="Description of good or service.")
quantity: float = Field(..., gt=0, description="Item quantity.")
unit_price: float = Field(..., ge=0, description="Unit cost.")
total_price: float = Field(..., ge=0, description="Line item total.")
class InvoiceDocument(BaseModel):
model_config = {"strict": True}
invoice_number: str = Field(..., description="Invoice unique identifier.")
vendor_name: str = Field(..., description="Supplier company name.")
invoice_date: str = Field(..., description="Date of issuance YYYY-MM-DD.")
currency: str = Field(default="USD", description="ISO 4217 currency code.")
line_items: List[LineItem] = Field(..., description="Line items.")
total_amount: float = Field(..., gt=0, description="Invoice grand total.")
@field_validator("currency")
@classmethod
def validate_currency(cls, value: str) -> str:
if len(value) != 3:
raise ValueError("Currency must be 3 uppercase letters.")
return value.upper()
def collect_textract_text(job_id: str) -> str:
lines: List[str] = []
next_token = None
while True:
params = {"JobId": job_id, "MaxResults": 1000}
if next_token:
params["NextToken"] = next_token
response = textract_client.get_document_text_detection(**params)
for block in response.get("Blocks", []):
if block.get("BlockType") == "LINE":
lines.append(block.get("Text", ""))
next_token = response.get("NextToken")
if not next_token:
break
return "\n".join(lines)
def run_extraction_pipeline(job_id: str, s3_uri: str) -> dict:
raw_text = collect_textract_text(job_id)
tool_spec = {
"tools": [{
"toolSpec": {
"name": "save_invoice",
"description": "Store verified invoice structure.",
"inputSchema": {"json": InvoiceDocument.model_json_schema()}
}
}],
"toolChoice": {"tool": {"name": "save_invoice"}}
}
converse_resp = bedrock_client.converse(
modelId="anthropic.claude-3-5-haiku-20241022-v1:0",
messages=[{"role": "user", "content": [{"text": f"Extract invoice fields from OCR:\n{raw_text}"}]}],
toolConfig=tool_spec,
inferenceConfig={"temperature": 0.0, "maxTokens": 2048}
)
blocks = converse_resp["output"]["message"]["content"]
tool_use = next(b["toolUse"] for b in blocks if "toolUse" in b)
validated = InvoiceDocument.model_validate(tool_use["input"])
db_item = {
"pk": f"VENDOR#{validated.vendor_name.upper().replace(' ', '_')}",
"sk": f"INVOICE#{validated.invoice_number}",
"vendor_name": validated.vendor_name,
"invoice_date": validated.invoice_date,
"currency": validated.currency,
"total_amount": str(validated.total_amount),
"line_items": [i.model_dump() for i in validated.line_items],
"source_s3_uri": s3_uri,
"job_id": job_id,
"status": "COMPLETED"
}
table.put_item(Item=db_item)
return {"status": "SUCCESS", "invoice_number": validated.invoice_number}
The data stays clean. Test this immediately against your raw files.
Common failure modes and operational traps
Production workloads expose subtle edge cases that local testing hides.
First, Step Functions has a 256 KB state payload ceiling. If your Lambda attempts to pass millions of raw Textract block characters directly through the state machine context, execution aborts with a States.DataLimitExceeded error. Write raw OCR text to an intermediate S3 key under an ephemeral prefix, passing only the resulting S3 pointer between states.
Second, Textract will occasionally see a coffee mug stain on a scan and parse it as a valid signature block. When processing degraded scans, OCR confidence drops. You should inspect average block confidence and reject pages whose confidence falls below seventy percent before invoking Bedrock.
Third, Claude 3.5 Haiku may occasionally produce line items whose sum does not match the printed total. You can enforce consistency with a Pydantic model validator:
from pydantic import model_validator
class ValidatedInvoice(InvoiceDocument):
@model_validator(mode="after")
def verify_totals(self) -> "ValidatedInvoice":
calculated = sum(item.total_price for item in self.line_items)
if abs(calculated - self.total_amount) > 0.05:
raise ValueError(f"Line items sum ({calculated}) mismatches total ({self.total_amount}).")
return self
Catching math errors before database insertion prevents poisoned accounting tables.
Limitations and operational trade-offs
Although DPIPE slashes operational costs by over 80 percent, my tests confirm two notable trade-offs.
Processing latency represents the first constraint. Textract asynchronous jobs take between 15 and 45 seconds to scan a ten-page document. Claude 3.5 Haiku requires an additional 2 to 4 seconds for semantic extraction. DPIPE suits asynchronous batch ingestion, not synchronous web requests where a user expects sub-second UI updates.
Document size introduces the second boundary. For massive PDF dossiers containing 200 pages, feeding the entire linearized OCR transcript in a single prompt exceeds reasonable context boundaries and increases latency. In our work with specialized engineering documents, we split massive documents into page batches or map states, a pattern detailed in our guide on multi-page PDF extraction with Textract.
This could easily be tested on complex legal contracts or multidimensional logistics manifests.
The wider perspective
More generally, modern Document AI is no longer a battle of fragile OCR bounding boxes against messy pixel alignments.
By decoupling character detection from semantic reasoning, you transform unstructured PDFs into verified software assets. Raw OCR handles pixels at rock-bottom commodity rates, while compact foundation models handle business logic with precision.
Once the pipeline runs quietly in production, your stakeholders will believe document extraction is pure magic, which is convenient because explaining OCR confidence scores at four in the afternoon is bad for mental health!
If you are designing resilient document agents or serverless cloud pipelines, explore our bespoke consulting services or read our latest technical papers on agentic AI architectures.