AI & TrustThe Blog
Before You Analyze a Contract, You Have to Know What It Is
A capable model may know what a contract is. That does not mean it will reliably know what deserves attention in every type of deal.
Contents9 sections
- One word was doing too much work
- The categories came from real agreements
- Thirty-seven doors instead of one
- Knowledge and attention are not the same thing
- The document itself should be the input
- Wrong category is now a real failure
- Structure is not the same thing as intelligence
- More deliberate does not mean proven
- The user is part of the classification system
From the WALDHORN.AI Build Archive
Build date: May 21, 2026
Someone can upload the right file and still get the wrong review.
The PDF can be perfectly legible. The language can be read correctly. The analysis can sound intelligent, detailed, and completely plausible.
The mistake can happen before any of that.
The system can be looking at the document through the wrong kind of agreement.
Review already exists inside WALDHORN.AI. It can take contract language, analyze it, and return a structured memo organized around risk. The problem I am running into is that the word contract hides far too much.
A recording agreement is a contract.
So is a publishing agreement.
So is a management agreement, producer agreement, synchronization license, distribution agreement, or 360 deal.
Treating all of them as one analytical object is convenient.
I do not think it is good enough.
One word was doing too much work
Different agreements can contain much of the same legal vocabulary.
Term. Territory. Ownership. Royalties. Approvals. Accounting. Termination. Warranties.
That overlap makes generic analysis tempting because a capable language model can usually discuss all of those concepts.
But recognizing a concept is not the same thing as knowing how much attention it deserves in a particular transaction.
A provision that is routine in one agreement can be economically central in another. The absence of a protection can be unremarkable in one deal and extremely important in another. Two clauses that look reasonable separately can become problematic because of how they interact inside a specific structure.
The useful question is not simply:
Is there something risky in this clause?
It is:
What should matter when this clause appears inside this kind of deal?
That is a much narrower problem.
I think narrower is better.
The categories came from real agreements
I did not want to invent thirty-seven categories and then write thirty-seven theoretical checklists from a blank page.
The analytical frameworks behind Review are being built from actual contracts.
I have gone through hundreds of real agreements across the kinds of music-industry deals WALDHORN.AI is meant to review. The useful part is not simply seeing how attorneys phrase the same clause differently.
It is seeing what repeatedly goes wrong.
An economic term that looks harmless until another section changes its meaning. A protection that is missing entirely. A rights grant that reaches further than the rest of the deal seems to suggest. An approval that exists but is narrower than it first appears. Two provisions that quietly work against each other.
Those patterns accumulate.
They also differ by agreement type.
The model already knows a great deal about contracts. I am not trying to teach a foundation model what a publishing agreement is.
I am trying to give the Review system a deliberate field of attention based on what has actually appeared in real agreements.
That is a different goal.
The model still has to interpret the document in front of it.
The system should make it harder for important classes of problems to depend entirely on what the model happened to notice on that run.
Thirty-seven doors instead of one
The new Review flow starts with 37 agreement categories.
The user chooses the type first, then uploads the document. That category determines the analytical framework used to examine it.
The result can still follow one consistent Review structure. The document is not forced through one universal idea of what a contract review should care about.
This adds friction.
One upload button would be easier.
A single option called Contract would be easier.
Asking somebody whether they are holding a recording agreement, publishing deal, management agreement, producer agreement, sync license, or another instrument makes the beginning of the workflow more demanding.
I am accepting that tradeoff.
The categories are not there so the product can advertise a large number.
They are there because the questions that should receive deliberate attention change with the transaction.
A recording agreement should be reviewed as a recording agreement.
A publishing agreement should be reviewed as a publishing agreement.
That sounds obvious when written down.
It becomes less obvious when the thing doing the analysis is a language model that is perfectly willing to answer almost any plausible question without first stopping to challenge the frame.
Knowledge and attention are not the same thing
This distinction is becoming important to me.
A strong model can know the answer to something and still fail to make that thing important in a particular response.
It has an enormous space of possible observations available to it.
Ask it to "review this contract," and it has to decide what deserves attention, how deeply to inspect each issue, what to omit, and how to distribute its limited output across the entire agreement.
That decision is probabilistic.
Two intelligent reviews can emphasize different things.
That is not necessarily a defect in the model. Human reviewers do this too.
But if I am building a product around repeatable contract analysis, I do not want the entire coverage of the review to be rediscovered from scratch every time.
The category framework gives the analysis a shape before interpretation starts.
Certain classes of questions belong in the field of attention because the underlying agreements have shown that they matter.
That does not force the conclusion.
It makes the coverage more deliberate.
The goal is not to make the model know more. It is to make the review depend less on what the model happened to notice.
The document itself should be the input
I also removed the paste-text workflow.
Review now accepts the actual document: PDF, JPEG, or PNG.
The uploaded file is stored and read directly through the multimodal analysis path. The current upload limit is 15 MB.
This solves a much less glamorous problem.
Copying a contract into a text box changes the thing being reviewed.
Page structure disappears. Tables can flatten. A scanned agreement becomes whatever another extraction step decided the text contained. Signature pages can lose their relationship to the provisions before them.
I do not want the user to prepare the contract for the machine.
The machine should meet the document closer to the form in which the person received it.
That does not make document understanding perfect.
A poor scan is still a poor scan. A model can still misunderstand language. A strangely constructed PDF can still be strange.
But an unnecessary transformation is an unnecessary opportunity to lose information.
Removing one is useful.
Wrong category is now a real failure
Once the category changes the analysis, the product has to be willing to admit when the category is wrong.
Otherwise the whole system is cosmetic.
If someone chooses one type of agreement and uploads something materially outside that category, WALDHORN.AI should not confidently finish the requested review just because the model can still read the words.
So category validity is now part of the Review flow.
A document can fail the selected category check.
When it does, the review stops.
The credits are returned.
The same refund behavior applies when the review itself fails.
I like this more than forcing every upload to produce something that looks valuable.
A paid AI product has an obvious incentive to always answer. The user uploaded a file. Computation started. A long analysis feels more complete than a refusal.
But a refusal can be more informative than a confident answer to the wrong question.
"Wrong type" says less.
It also claims less.
Structure is not the same thing as intelligence
It would be easy to describe what I built as better instructions for a model.
Technically, instructions are part of the mechanism.
That description misses the interesting part.
The work is deciding what deserves reliable attention before the model starts interpreting a particular document.
The frameworks are an attempt to turn observations from real agreements into repeatable analytical coverage.
The model still decides what the language means.
The framework decides what kinds of questions should not be left entirely to chance.
I think those responsibilities should stay separate.
If a producer agreement repeatedly creates a particular class of economic problem, I do not want the system to rediscover the importance of that class only when the model happens to focus on it.
If a management agreement makes another group of provisions especially consequential, that should affect what the review deliberately examines.
This is less about making AI more knowledgeable than making the product more disciplined about how that knowledge is used.
More deliberate does not mean proven
There is an uncomfortable limit to everything I built here.
A more structured review is not automatically a better review.
I can decide what the system should examine.
I can make the analytical coverage more repeatable.
I can narrow the context to a specific kind of agreement.
I can reject a document that does not belong in that context.
None of those things prove that the conclusions are correct.
A beautifully structured review can still contain a bad judgment.
That is a different problem.
I do not have a complete answer to it yet.
There is an important line somewhere between controlling what the system looks at and proving that what it says about those things deserves confidence.
For now, I am working on the first side of that line.
The user is part of the classification system
There is also an obvious weakness in the workflow I built.
If correct classification matters, I am currently asking the user to do the classification.
Someone who knows exactly what they have is fine.
Someone holding a twelve-page PDF titled only Agreement may not be.
A musician should not need to become good at legal taxonomy before a contract-review product can help them.
The system can now recognize that the selected category does not fit the uploaded document.
That is useful.
But it creates a slightly absurd interaction:
This is the wrong category.
Okay.
Which category is it?
WALDHORN.AI cannot answer that reliably yet.
So I have made the review stricter and, at the same time, exposed the next problem in the workflow.
I am fine with that.
A good boundary often reveals the problem that was hiding behind it.
Before WALDHORN.AI analyzes what a contract says, I want the system to have a defensible answer to a simpler question.
What kind of document is this?