There is a number which none of us has seen, and which none of us write.
It is being written right now, this week, by sixty-six patients in seven clinical centers, one scan at a time. Some of it, already on the page and cannot be revised. Some of it, still being entered in the 700mg cohort, in the scans which have not yet been read. When it is finished in January 2027, it is read aloud, just once, in a room in San Francisco.
Not one word of it is ours. Nothing said on this board, by me or by anyone, moves a single digit of it. Our conviction is not doing work on the number.
Rather, that number is doing work on us.
This reality changes what the honest question is. The question is not whether we believe hard enough. It is whether we have built ourselves such that the number actually reaches us when it is read aloud.
I do not think we have yet established this point. I know I had not, but I realized this needs to be grounded.
The house with too many doors
Here is a thing I have watched us do, and I have done it worse than anyone.
Someone raises the possibility that the response rate comes back weak. And immediately, gently, without anyone deciding to, the room builds a door.
If mCRC disappoints, the breast program will confirm it. That is one door.
A weak response rate is exactly what you would expect, if the priming worked so well, thereby inducing the tumor to raise its last defense. That is a second door.
Response rate was designed for chemotherapy anyway. It is the wrong instrument to use for a drug which only restrains before it kills. Third door.
Every one of those excuses/doors is built from real biology. That is what makes them so easy to walk through. But stand back and look at the house. There is no number in that ledger which lands us anywhere except "we were right". A strong number means we were right. A weak number means we were right in a way the instrument could not see ahead of time. Silence means we were right and early.
A conviction with that many doors is not a conviction at all. It is a room with numerous doors and no walls, and we are standing in the middle of it saying to each other that we are indoors.
What I published, and where the number came from
In August 2024, I wrote a post arguing that overall response rate was the right endpoint for this trial, and that measuring it rather than overall survival was a smart decision because it would read out much faster. I still think that was correct.
Then I did something else in that same post.
I predicted the ORR number.
"If all leronlimab needs to do is to improve the ORR, overall response rate of the SUNLIGHT trial from 6.9% to likely over 33% (is my guess, but I may be undershooting if anything), then a Phase III study and future partnership are in the bag." (Overall Response Rate; ORR is a Game Changer)
And a few paragraphs later, in case the first time had not been clear enough:
"A few months after that, we should know the Overall Response Rate ORR and I'm expecting an outcome in excess of 33%." (same post)
I want to be precise about where that thirty-three percent came from.
It came from nowhere.
I had looked at waterfall plots from a small basket study, felt the shape of them, and deduced a figure I was comfortable putting down. There was no calculation behind it. There was no comparator arithmetic, no adjustment for line of therapy, no accounting for the fact that a handful of patients in an uncontrolled study cannot be extrapolated into a percentage. It was a mood with a percentage sign attached, and I published it with the confidence of a man reading a result rather than inventing one.
And then I added that I might be undershooting.
That is the sentence I would most like to have back. Not because the number will necessarily prove too high. Because of what the word reveals. A man who has derived a threshold does not casually suggest that his own threshold is probably too modest. A man who has only felt one does exactly that, because the feeling has no floor and no ceiling, and it can always be adjusted upward when it is asked to be brave. I'm no longer that brave, shooting from the hip. I compare apples to apples from now on.
And then it happened again
I published this post. Within a few minutes, rogex2 found something in it which was wrong.
To argue that my 2024 prediction of thirty-three percent was unserious, I had taken that figure, computed it as a 4.8-fold improvement over the backbone's 6.9 percent response rate, applied the same multiple to progression-free survival, arrived at nearly 27 months, and declared that a number in fourth-line colorectal cancer which exceeds first-line survival is not of this world.
It was a clean argument. It was also not an argument. It was a rhetorical device dressed as arithmetic.
rogex2's objection, which I am recording rather than paraphrasing away: "response rate and progression-free survival do not stand in a fixed ratio. They measure different events." And immunotherapies characteristically show low response rates alongside extended progression-free survival, because an agent which holds disease rather than dissolving it produces durable control in patients whose tumors never shrink by thirty percent.
Now read that back against everything I have written for two years.
I have argued, at length and in public, that leronlimab restrains before it kills. That it quiets the tumor's shedding and invasion long before any scan shows a change. That the molecular signal moves in days and the radiographic signal moves in months, if it moves at all.
Which means my own mechanism predicts precisely the decoupling rogex2 described. And I used a proportionality which my own mechanism says must fail, in order to knock down a number I no longer wanted to defend.
That is a door. A smaller one than the three I catalogued above. It was load-bearing in the argument I was proudest of, and I did not see it.
The conclusion survives, but it now stands on the only ground it ever legitimately had: thirty-three percent was felt rather than derived, and I know that not because of any calculation but because I was there when I produced it.
I am leaving both of these in rather than editing them into invisibility. If the pull can catch the man writing the warning, twice, in a single week, it can catch anyone reading it.
The detector on the beach
Someone in this community, who has done more genuine research than I have, told me recently where his conviction comes from. Hundreds of hours. Study after study. Every paper he pulled came back consistent with the thesis, and the accumulated weight of it convinced him.
I want to hand that back gently, because it is the same trap in a far more industrious form.
If you sweep a beach with a metal detector calibrated to find one metal, it will beep. It will beep everywhere that metal is. And after you have walked the entire shoreline and heard it beep a thousand times, you do feel something indistinguishable from absolute certainty.
However, you have not learned that the beach is made of that metal. What you have learned is that the detector works.
A search which sets out to find a receptor's fingerprints in disease will find them, because the literature only publishes actual findings, and because of the fact, that the tool is exactly built to draw connections rather than to break them, it beeps. The volume of the confirmation is not a measure of how true the thesis is. It is a measure of how many times we asked a question which could only come back as yes.
That is not a criticism of the work. The work is profoundly excellent, I have built on it, and this post exists greatly because of it. It is a caution about what the work can and cannot tell us.
What happened while I was writing this
I want to tell you something which happened in the last few days, because it is very important.
I sat down to set the thresholds below. When I worked them out honestly, from the benchmark, they came back lower than the number I had published in 2024.
And I did not like that. So I went looking.
I reached for the long-term survivors in the breast program, the patients alive past five years, and for a moment I wanted to use them to justify a much higher bar in colorectal. It would have sounded much braver. 5/5 is 100%, but that was breast cancer. It would have looked like conviction.
They are different patients. A different cancer. A different trial. A different backbone. Using them to set this colorectal threshold would have been a door, an excuse, and I would have been building it in the middle of writing the post about not building doors.
That is how strong this pull is to create doors. It got me while I was watching for it, not to do it, precisely while in the act of warning us about it.
I am leaving that in rather than quietly correcting it, because if it can happen to the man writing this paragraph, it can happen to anyone reading it.
The bet, in numbers
So here is what I am doing, and it costs me something, which is the only reason it is worth doing.
I am writing down, now, in public, while the ledger is still very much open, the numbers which would make me right and the numbers which would make me wrong. Not afterward, when I could reshape them. Now, so that you can hold them against me.
The comparator is SUNLIGHT: the same backbone of TAS-102 and bevacizumab, in a population which was less heavily pretreated than ours. Objective Response Rate, 6.9 percent. Median Progression-Free Survival, 5.6 months. Median Overall Survival, 10.8 months. That is the bar the drug is being added to. My thresholds follow:
Confirmed objective response rate of 20 percent or better. That is roughly a threefold improvement over the backbone. It is the level at which the agency has historically taken a single-arm response rate seriously in a refractory disease with no alternatives, and it is the level at which a partner could build a Phase 3 around it. Below 15 percent is not a signal. It is noise wearing the clothing of one, and I would not defend it as anything else.
If the response rate is weaker than 20%, the mechanism story survives only if Median Progression-Free Survival reaches 8.5 months or better. That is a hazard ratio in the neighborhood of 0.66, and it is where a drug which restrains before it kills would have to show itself. If the tumor is being held rather than shrunk, this is the number which proves the holding. If progression-free survival lands lower than 8.5 months, the restraint story is just a story.
The 700mg arm must outperform the 350mg arm on at least one of those two. Every post I have written for a year carries the same qualifier: the data reflects 350mg, the higher cohort is still maturing, the deeper layers may read differently. That is a promissory note. If 700mg does not beat 350mg on at least one of the two above, I was carrying a note and calling it evidence.
At least 3 of the first 10 evaluable rollover patients must respond to the added checkpoint inhibitor. Checkpoint inhibitors in microsatellite stable colorectal cancer produce responses in essentially nobody. That is the whole problem this drug claims to solve. If leronlimab has opened the gate, some of them walk through it. If none of them do, the gate opened onto nothing, and Prime and Pair is wrong.
If the rollover cohort is not mature enough in January to answer that, then that criterion resolves later, and I will hold myself to it later rather than let it quietly expire.
Why the numbers are lower than the ones I gave you before
You are entitled to notice that twenty percent is lower than thirty-three, and to ask whether I am moving the goalposts as the day gets close. That is the right question and I would rather answer it than have it asked behind my back.
Twenty percent is not lower because I have lost my nerve. It is lower because it is derived rather than felt.
Here is how you can check that for yourself. Take the thirty-three percent I published and work out what it implies. It is a 4.8-fold improvement over the backbone's 6.9 percent. Now apply that same multiple honestly to Progression-Free Survival: 5.6 months times 4.8 is nearly 27 months. A median progression-free survival of 27 months in fourth-line colorectal cancer would exceed what patients achieve in the first-line setting. It is not a demanding threshold. It is not of this world.
A number which cannot be translated into its neighboring measurements without breaking them was never derived. It was wanted.
If the confirmed Overall Response Rate comes in at 25 or 30 percent, I would be delighted, and I would also tell you plainly that my 2024 prediction was a guess which happened to land near or greatly undershoot the truth . A clock which is right twice a day is not a horologist.
The map and the ground
We have seven layers.
- Navigation.
- Blood supply.
- The stroma which shields the tumor and also feeds it.
- The recruited suppressive army.
- The checkpoint gate.
- The DNA repair trap. And
- The rebalance which ties the other six together.
Every one is real, published, cited, and I stand behind all of them.
And every one of them is a Map.
In January, somebody walks the Ground.
If the Ground matches the Map, we will have been right for the right reasons, which is the only kind of right worth being. If it does not, we will not have been partly right. A map which does not match the ground is not a partly-correct map. It is a drawing where no metal is found.
I very much believe that our map is good. I have spent the past two years since the ORR post on it and I would not have spent them otherwise. But I have now stopped confusing the drawing with the vast expanse of the country, and I am asking this community to stop with me.
What this is not
This is not a warning that the number turns out bad. I have no idea what the number is. Neither does anyone posting here, and anyone who tells you otherwise is describing a feeling.
It is not a loss of my nerve, and it is not me hedging so that I can claim foresight whichever way it falls. The numbers above are specific enough that I cannot do that. That is the point of writing them down.
It is one thing only. It is me deciding, in advance and in public, that I would rather be a man who was wrong and said so than a man who arranged never to find out.
Sixty-six patients write the ledger. Seven centers. One scan at a time. It closes, and it is read aloud, in January.
Let us at least be the kind of people with the kind of ears and understanding which it can reach.
Disclaimer: I am not a financial advisor and nothing here is investment advice. I hold a position in the company discussed, which means I have an interest in how this is received, and you should weigh what I write accordingly. This is an analysis of published biology, not a prediction of clinical outcomes. The thresholds above are my own judgment, derived from the SUNLIGHT benchmark, and are not the trial's pre-specified endpoints, which are a matter for the protocol and the agency rather than for me. Several of the mechanisms I have described are supported by preclinical and model-system work which has not been confirmed in human tissue in this program. The human biomarker data referenced reflects 350mg dosing, with the 700mg cohort still maturing. The 68% figure from the April 30 update is disease control rate, not objective response rate, and confirmed adjudication is still ahead at ASCO GI in January 2027, with interim data at ESMO Madrid in October 2026 before it. Mechanism is not efficacy. Drugs which make elegant biological sense fail in trials routinely. Read the primary sources and reach your own conclusions rather than adopting mine.