Sep 2026 · 12 min read
PublishedIt Parsed. It Was Wrong Anyway.
I built a small contract-extraction pipeline that appeared to work perfectly. The model returned a valid Pydantic object. Every required field was present. The dates were valid Python date objects. Nothing crashed.
The data was still wrong.
One contract said September 20, 2026. The pipeline returned 9992-09-20. Another said August 14, 2026 and came back as 8358-08-14.
That failure changed how I think about evaluating LLM systems. A successful parse only proves that the model followed the required structure. It does not prove that the extracted information is correct.
This post is both the story of that bug and a lab you can run yourself. While making this post, I am rerunning each stage and adding the exact outputs before I publish it. Doing that work again is part of the point: I want to understand the system, not just preserve a result I got once.
12 sections
01What you will build
A pipeline that extracts six fields from unstructured service contracts: client name, agreement date, service, project amount, deposit required, and completion date.
Then you will test it at three levels:
- 01Parse success: did the API return a valid object?
- 02Field accuracy: did each extracted value match an answer key?
- 03Business validity: can a value pass type validation and still be impossible for the application?
The full version of my lab used ten synthetic contracts. The compact version below uses three representative cases so the important parts are easy to follow: a clean contract, a percentage deposit, and a contract with outdated values that must be ignored.
02Setup
Everything you need is in this post — copy each block in order into a file or notebook of your own and it runs top to bottom. Before you start you will need Python 3.10 or newer, and an OpenAI API key on an account with billing enabled.
Following the compact version costs nine API calls: one baseline, one after the schema fix, three for the evaluation, two for the model comparison, and two for the grounding test. They are small requests against a cheap model, but they are not free — check your own pricing page before you start.
I used Python, the OpenAI SDK, Pydantic, and python-dotenv. With uv, create a project and install the dependencies:
uv init contract-extraction-labcd contract-extraction-labuv add openai pydantic python-dotenv ipykernelThat creates a .venv folder inside the project. ipykernel is only needed if you plan to work in a notebook — but installing it now saves a detour later, and it is the piece people most often discover they are missing.
Script or notebook — pick one
Every code block below works either way. The script path has fewer moving parts:
# Paste the blocks into lab.py as you go, then run it:uv run python lab.pyuv run uses the project environment automatically, so there is nothing to activate and no kernel to choose. If you would rather work cell by cell, use a notebook — which does need one extra step.
Selecting the kernel in a notebook
A new notebook does not automatically use the environment uv just built. If you skip this, your imports fail with ModuleNotFoundError even though the packages are installed — the notebook is running against a different Python.
In VS Code: create the notebook inside the project folder, click Select Kernel at the top right, choose Python Environments, and pick the entry pointing at .venv in your project directory. Not your system Python, and not a global conda environment.
If you prefer Jupyter in the browser, uv can launch it against the project environment in one command, and the kernel will already be correct:
uv run --with jupyter jupyter labTo confirm you are on the right interpreter before going further, run this in the first cell. The path it prints should be inside your project's .venv:
import sysprint(sys.executable)Store the API key in a local .env file and keep that file out of version control:
OPENAI_API_KEY=your-key-hereStart the notebook or Python file with:
from datetime import datefrom typing import Annotatedimport re
from dotenv import load_dotenvfrom openai import OpenAIfrom pydantic import BaseModel, Field, WithJsonSchema, field_validator
load_dotenv(override=True)openai = OpenAI()RERUN CHECKPOINT — confirm the environment loads and the client initializes without printing or committing the key.
03Step 1: Start with the obvious schema
My first schema used Python date fields directly:
class Contract(BaseModel): client_name: str agreement_date: date service: str project_amount: float deposit_required: float completion_date: dateUse one clean contract as a baseline:
synthetic_contract = """SERVICE AGREEMENT
Client: Acme Plumbing LLCAgreement Date: September 20, 2026Service: Website redesign and lead-generation setupProject Amount: $4,800Deposit Required: $2,400Completion Date: November 15, 2026"""
system_prompt = """You are a contract data extraction assistant.Extract information only from the provided contract.Do not invent or infer information that is not present."""
response = openai.responses.parse( model="gpt-5.4-nano", input=[ {"role": "system", "content": system_prompt}, {"role": "user", "content": synthetic_contract}, ], text_format=Contract,)
print(response.output_parsed)My original run produced a valid object with a corrupted year. Because model behavior can change, your rerun may or may not reproduce the same failure. Record the exact model name, date, output, and generated schema instead of forcing the old result.
print(Contract.model_json_schema()["properties"]["agreement_date"])RERUN CHECKPOINT — paste the baseline object and the agreement_date schema fragment here.
04Step 2: Change what the model sees
The bad values were valid calendar dates, so Pydantic accepted them. Type validation had done its job: the value was a date. The business meaning was wrong.
The LLM-facing JSON schema was the important clue. A Pydantic date added "format": "date". In this setup, that format constraint was associated with the corrupted generations. I changed the model-facing schema to a plain string while preserving Python date parsing and validation:
LLMDate = Annotated[ date, WithJsonSchema({"type": "string"}),]
DATE_HINT = ( "Return as YYYY-MM-DD, e.g. 2026-08-14. Slash dates like 9/2/26 are " "US month/day/year, and a 2-digit year like 26 means 2026.")
class Contract(BaseModel): client_name: str = Field( description="The business name only: the company doing the work." ) agreement_date: LLMDate = Field( description=f"Date the agreement was made. {DATE_HINT}" ) service: str = Field(description="Short description of the work.") project_amount: float = Field( description="Final total price in dollars. If the price changed, " "use the final agreed price." ) deposit_required: float = Field( description="Deposit in dollars. If given as a percentage, calculate " "the dollar amount. Use 0 if no deposit is required." ) completion_date: LLMDate = Field( description=f"Final completion date. {DATE_HINT}" )
@field_validator("agreement_date", "completion_date") @classmethod def year_must_be_realistic(cls, value: date) -> date: if not 2000 <= value.year <= 2100: raise ValueError(f"Unrealistic year {value.year} in {value}") return valueThis creates two layers of protection. Prevent: show the model a plain string field with an explicit YYYY-MM-DD instruction. Catch: convert the string to a Python date and reject years outside the application's allowed range.
Schema descriptions guide the model. Validators enforce business rules after the response arrives.
Run the clean contract again with the revised model, then inspect the value and its Python type:
response = openai.responses.parse( model="gpt-5.4-nano", input=[ {"role": "system", "content": system_prompt}, {"role": "user", "content": synthetic_contract}, ], text_format=Contract,)
result = response.output_parsedprint(result)print(type(result.agreement_date))RERUN CHECKPOINT — paste the post-fix object and date type. If the baseline no longer fails, say so honestly and keep the validator as a regression guard rather than claiming the fix caused the newer result.
05Step 3: Build a compact test set
A clean document can make a weak system look reliable. These three inputs increase the difficulty without making the article unreadable.
contract_1 = """Client: Boston Roofing Inc.Agreement Date: August 14, 2026Service: Roof replacementProject Amount: $12,750Deposit Required: $5,000Completion Date: October 1, 2026"""
contract_2 = """Subject: Re: Kitchen remodel - confirming details
Harbor Kitchen & Bath will handle the full kitchen remodel.The total comes to $18,400. We ask for a 30% deposit.We signed off on June 9, 2026 and will finish by August 21, 2026."""
contract_3 = """Cape Cod Decking originally quoted $15,000 for a new composite deck.After the customer chose a smaller layout, the final agreed price wasreduced to $13,200. The deposit was first set at $3,000 but was loweredto $2,500. The agreement was signed on September 1, 2026. The originalcompletion date of November 1, 2026 was pushed back, and the work mustnow be finished by November 20, 2026."""
contracts = [contract_1, contract_2, contract_3]Write the expected values before running the evaluator. This prevents you from changing the answer key to match the model after you see its output.
expected_results = [ { "client_name": "Boston Roofing Inc.", "agreement_date": "2026-08-14", "service": "Roof replacement", "project_amount": 12750, "deposit_required": 5000, "completion_date": "2026-10-01", }, { "client_name": "Harbor Kitchen & Bath", "agreement_date": "2026-06-09", "service": "Kitchen remodel", "project_amount": 18400, "deposit_required": 5520, "completion_date": "2026-08-21", }, { "client_name": "Cape Cod Decking", "agreement_date": "2026-09-01", "service": "Composite deck", "project_amount": 13200, "deposit_required": 2500, "completion_date": "2026-11-20", },]
assert len(contracts) == len(expected_results)The second contract checks whether the model can calculate a dollar deposit from a percentage. The third checks whether it can choose final values instead of earlier ones.
My full ten-contract suite also included narrative prose, a casual text message, formal legal language, a pipe-delimited table, rushed notes, a no-deposit case, and a French-language contract.
06Step 4: Test the scoring logic first
A broken grader can produce a confident but wrong score. I used comparison rules that fit each data type:
SCORED_FIELDS = [ "client_name", "agreement_date", "project_amount", "deposit_required", "completion_date",]
def normalize(text: str) -> str: """Lowercase, remove punctuation, and collapse spaces.""" text = text.lower() text = re.sub(r"[^\w\s]", "", text) return " ".join(text.split())
def is_match(field: str, expected, actual) -> bool: """Compare one field using a rule that fits its type.""" if field in ("project_amount", "deposit_required"): return abs(float(actual) - float(expected)) < 0.01
if field == "client_name": options = expected if isinstance(expected, list) else [expected] return normalize(actual) in [normalize(option) for option in options]
if field in ("agreement_date", "completion_date"): if isinstance(actual, date): actual = actual.isoformat() return actual == expected
return actual == expected
assert is_match("project_amount", 5520, 5520.0000001)assert is_match( "client_name", "Northeast Electrical Contractors, Inc.", "northeast electrical contractors inc",)assert is_match("completion_date", "2026-10-01", date(2026, 10, 1))assert not is_match("deposit_required", 2500, 3000)print("Scoring helpers passed")I left service out of the automatic score. Exact string matching is a poor measure when two descriptions can have the same meaning with different wording. In this small lab, I review that field manually.
RERUN CHECKPOINT — keep the helper assertions in the published code and paste the exact confirmation line from your run.
07Step 5: Measure parsing and accuracy separately
The evaluator counts a failed parse as five incorrect fields. A failed document cannot disappear from the denominator just because there is no object to compare.
MODEL = "gpt-5.4-nano"best_prompt = """You are a precise contract extraction system.
Read the contract carefully and extract the requested structured fields.
Rules:- Only use information explicitly written in the contract.- Never fabricate values.- Do not confuse the project amount with the deposit amount.- Preserve the correct client, dates, service, and monetary amounts."""
parsed_ok = 0correct_fields = 0total_fields = 0failures = []service_reviews = []
for index, (contract_text, expected) in enumerate( zip(contracts, expected_results), start=1): try: response = openai.responses.parse( model=MODEL, input=[ {"role": "system", "content": best_prompt}, {"role": "user", "content": contract_text}, ], text_format=Contract, ) result = response.output_parsed except Exception as error: print(f"Contract {index}: API or parse error: {error}") result = None
if result is None: total_fields += len(SCORED_FIELDS) failures.append((index, "ALL", "parsed Contract", None)) continue
parsed_ok += 1 actual = result.model_dump()
for field in SCORED_FIELDS: total_fields += 1 if is_match(field, expected[field], actual[field]): correct_fields += 1 else: failures.append((index, field, expected[field], actual[field]))
service_reviews.append((index, expected["service"], actual["service"]))parse_rate = parsed_ok / len(contracts) * 100field_accuracy = correct_fields / total_fields * 100
print(f"Parse success: {parsed_ok}/{len(contracts)} = {parse_rate:.0f}%")print( f"Field accuracy: {correct_fields}/{total_fields} " f"= {field_accuracy:.1f}%")print("Failures:")for failure in failures: print(failure)
print("Service review:")for review in service_reviews: print(review)RERUN CHECKPOINT — paste the exact printed summary including any failures, then explain the first failure in plain language. Do not reuse the earlier score unless the fresh run produces it.
In my original full evaluation, the fixed schema parsed all ten contracts and matched all 50 automatically scored fields. That was encouraging, but it was not proof of production reliability. The sample was small, synthetic, and partly written by the same person who wrote the answer key.
The manual service review also exposed limitations that the numeric score hid:
- 01The model preserved a typo from rushed notes: central ac instalation.
- 02One description absorbed distractor context about an earlier price and a smaller layout.
- 03The French contract stayed in French, while my answer key expected an English translation.
That last case was an evaluator problem, not necessarily a model problem. Ground truth — the answer key used to score a system — also needs review.
08Step 6: Compare cost only after quality
I compared two models on the same difficult contract. The important order is quality first, cost second.
def estimated_cost(usage, input_price, output_price): return ( usage.input_tokens / 1_000_000 * input_price + usage.output_tokens / 1_000_000 * output_price )
def run_model(model, input_price, output_price): response = openai.responses.parse( model=model, input=[ {"role": "system", "content": best_prompt}, {"role": "user", "content": contract_3}, ], text_format=Contract, )
cost = estimated_cost(response.usage, input_price, output_price) print(model) print(f"Input tokens: {response.usage.input_tokens}") print(f"Output tokens: {response.usage.output_tokens}") print(f"Estimated cost per contract: ${cost:.6f}") print(f"Estimated cost per 10,000: ${cost * 10_000:.2f}") print(response.output_parsed)Pricing changes, so look the rates up at the moment you run this rather than trusting a number copied from someone else's notebook. Leaving the placeholders as None is deliberate — the call fails loudly instead of quietly reporting a wrong cost:
# Look up the current per-1M-token prices first. Price both models from the# same source on the same day, or the comparison is not a comparison.NANO_INPUT, NANO_OUTPUT = None, NoneMINI_INPUT, MINI_OUTPUT = None, None
run_model("gpt-5.4-nano", NANO_INPUT, NANO_OUTPUT)run_model("gpt-5.4-mini", MINI_INPUT, MINI_OUTPUT)My earlier single-contract comparison found that both models were correct on the automatically scored fields, while the larger model cost more. One contract is not a benchmark. The useful lesson is the decision process: define the quality threshold, evaluate models against the same cases, and choose the least expensive model that meets the requirement.
RERUN CHECKPOINT — paste the current pricing source, exact token counts, estimated costs, and both model outputs. If the models differ in quality, discuss that before discussing price.
09Step 7: Make missing information predictable
I also removed the client's address from a contract and asked for it:
hallucination_contract = """Client: Acme Plumbing LLCAgreement Date: September 20, 2026Project Amount: $4,800"""
question = ( "What is the mailing address of the client in this contract? " "Give me the address.")
grounded_prompt = """Answer using only facts explicitly stated in the provided contract.If the requested information does not appear in the contract, respond with:NOT PROVIDEDDo not guess."""
response_a = openai.responses.create( model="gpt-5.4-nano", input=[ {"role": "user", "content": hallucination_contract}, {"role": "user", "content": question}, ],)
response_b = openai.responses.create( model="gpt-5.4-nano", input=[ {"role": "system", "content": grounded_prompt}, {"role": "user", "content": hallucination_contract}, {"role": "user", "content": question}, ],)
print("Without grounding:")print(response_a.output_text)print("With grounding:")print(response_b.output_text)In my earlier run, the ungrounded model did not invent an address. It explained that the address was missing. The grounded version returned exactly NOT PROVIDED.
That means I cannot honestly claim the prompt prevented a hallucination. What it improved was consistency: the missing value became a predictable, machine-readable response instead of a paragraph.
RERUN CHECKPOINT — paste both responses exactly as returned. If the ungrounded response changes, report what happened without rewriting the experiment to fit the old conclusion.
10Use the rerun as a learning exercise
Before looking at the old outputs, I will do four things:
- 01Predict: write down which contract I expect to fail and why.
- 02Explain: describe the difference between parsing, accuracy, and business validation from memory.
- 03Run: execute the notebook from a clean kernel in order, without skipping cells.
- 04Challenge: add one new adversarial contract, write its answer key first, and run the full evaluator again.
Good adversarial additions include an ambiguous slash date, a missing deposit, a negative amount, two companies in one document, or a completion date earlier than the agreement date.
The last example exposes an important next step. A year-range check is useful, but it does not prove that two individually valid dates make sense together. A production model could also validate that completion_date is on or after agreement_date.
11What I would carry into production
- 01Define every field before extraction. Client was initially ambiguous: did it mean the business doing the work or the customer receiving it? If the team cannot define a field, the model cannot apply the definition consistently.
- 02Keep parse success and accuracy separate. A valid object can contain incorrect values.
- 03Add business validators. Types catch malformed data; domain rules catch values that are structurally valid but impossible for the application.
- 04Count failures honestly. If a document does not parse, its fields still count against accuracy.
- 05Test the evaluator. Normalization and scoring logic need their own assertions.
- 06Review free text differently. Exact string matching is a poor measure for semantically equivalent descriptions.
- 07Compare cost only after quality. A cheap wrong answer is expensive once it enters a workflow.
- 08Keep adversarial cases. Short dates, percentages, conflicting values, abbreviations, missing fields, and multilingual text belong in the regression suite.
12The takeaway
Parse success tells me the model followed the format. Field accuracy tells me whether it was right.
The most useful part of the exercise was not getting a perfect score after the fix. It was watching a pipeline return polished, typed, completely believable wrong data — and then building the checks that exposed it.
Rerunning the code before I publish this is part of the lesson. The output belongs in the post only after I can reproduce it, explain it, and say exactly what the test does not prove.