Jamie Stanton

Summarization is the least valuable thing an IR AI can do

The first thing almost every IR team does with a new AI tool is paste in their own earnings release.

I've done it. It's a completely natural instinct. You have a document in front of you, the box is right there, and you want to know whether the thing is any good before you trust it with anything that matters. So you feed it something you know cold, and you read what comes back, and you check it against your own understanding.

And that is exactly the problem. You picked the one document in the world where a summary can teach you nothing, because you wrote it. You argued with legal about three words in the second paragraph. You know which number the CFO wanted bolded and why. When the summary comes back accurate, you've learned that the tool can read. You have learned nothing about whether it can think, and you have spent your evaluation on the single task where those two things are hardest to tell apart.

Verifiability and value run in opposite directions

The reason summarization dominates every demo, every pilot and every first week of usage is that it's the easiest output to check. You can look at a summary of your own transcript and grade it in ninety seconds. That feels like diligence.

Now try to grade an inference about why a top-fifteen holder trimmed. You can't, not immediately, maybe not for two quarters. The output you can verify fastest is the output that carries the least information, and the output that carries the most information is the one you cannot verify at all in the moment. Every evaluation process I've watched runs straight into this and most of them don't notice.

The downstream cost is worse than the wasted week. If your team's working standard for these systems becomes "did the summary read well," you will systematically buy the tools that write beautifully and reason poorly. Fluency becomes the purchase criterion. Fluency is free now. It has been free for a while.

Compression throws away the part you need

Set aside who wrote the document. Summarization is compression, and compression works by keeping what's central and discarding what's unusual. That is the precise opposite of what an IR professional needs from a document.

Think about what you were actually doing back when transcripts arrived as a paid PDF a day and a half after the call and you read them with a highlighter. You were not extracting the gist. You knew the gist. You were hunting for the delta. Which analyst asked a question nobody asked last quarter. Whether the CFO said "timing" this time where he said "headwind" before. The hedge that appeared in the guidance language between the draft you saw Tuesday and the words that came out of his mouth Thursday. The three-second pause.

None of that survives compression. All of it is, in the statistical sense, noise. It's also the entire job. A summary of a peer's earnings call tells you what the peer said, which you could have guessed. What you want to know is where they said something different from last quarter, different from you, or different from what the sell side expected them to say. That's not a summary. That's a comparison, and it requires the model to hold two documents at once and care about the gap between them.

The documents worth reading are the ones you didn't write

Here's the honest inventory. The material an IR team feeds these systems most often is the material it produced itself: the release, the script, the deck, the transcript of the call it hosted, the FAQ it drafted. Information gain on all of it is approximately zero.

The material with actual information in it is the material nobody on your team has time to read. Eleven peer transcripts. The sell-side note that quietly disagrees with your framing and never says so out loud. Quarterly commentary letters from your top holders, where a portfolio manager explains her process to her own investors in plainer language than she will ever use with you. Proxy advisor reports. The exhibits attached to a 13D. Conference Q&A from a session you weren't in the room for.

That's the asymmetric pile, and it's asymmetric precisely because it went unread. But notice that it doesn't want summarizing either. You don't need the gist of eleven peer calls. You need to know which two of the eleven said something that undermines the guidance you're about to give.

A ladder worth using

Roughly in order of how much value I've seen each rung produce, and roughly in inverse order of how well each one demos:

Compression. Tell me what this says. Near zero, for the reasons above.

Extraction. Pull the same fifteen fields out of forty documents and put them in a table. Deeply unglamorous, genuinely useful, and the rung most teams skip on their way to something more exciting. Structured beats eloquent.

Comparison. How does this document differ from the last one, from the peer's version, from what we said in March? This is where the value starts and it is where most tools quietly stop being able to help you, because comparison requires the system to actually have your history rather than just your prompt.

Contradiction. Where does this disagree with something else we believe, including something we've said publicly? This is the rung I'd pay for. It's also the one that requires a system willing to tell you that you are wrong, which is a product decision more than a technical one.

Interruption. Nobody asked, and it told you anyway, on a Tuesday, because something changed. Impossible to demo, hard to build, and the only rung that solves the problem you actually have, which is that the thing you needed to know surfaced on a day you were busy.

Where summarization does earn its keep

I want to be fair to it, because there's a real use and it gets confused with the fake one.

Summarization is valuable when the reader is not you. A board pre-read where four directors will not read eleven peer transcripts, and shouldn't have to. A new CFO who needs ten years of narrative arc in twenty minutes. A colleague in another function who needs the gist of a call they will never listen to. Getting the substance of something into the head of someone who was never going to read the source is a real job and these tools are good at it.

But that's distribution, not analysis. Summarization has genuine value as a distribution tool and almost none as an analytical one. Most of the disappointment I hear about AI in IR comes from teams who bought it for the second thing and are using it for the first.

An exercise, if you want to test this

Take last quarter's earnings call. Ask whatever tool you have to summarize it. Read the summary and count the things in it you didn't already know. Be honest with the count.

Then ask it something else: list every place the language in this quarter's answers differed from the way we answered the same question last quarter. Watch what happens. Most tools cannot do the second thing at all, and the ones that can will hand you something you'd have wanted on your desk the morning after the call.

That gap between the two answers is the whole argument. It's also, I'd guess, most of the value that's actually available to an IR team right now.

Written in a personal capacity. The views here are my own.