This is not betting advice. It is a public research journal: what we built, why we believed in it, where we were wrong, and what changed after we stopped treating a beautiful report as a passport to the future.
There are two ways to tell this story.
The heroic version says that we spent roughly a year and a half studying bookmaker markets, built complex pipelines, generated ensembles, tested millions of combinations, published sealed signals, hashed the files, designed a timestamp-proof layer, and nearly assembled a spacecraft.
The honest version says that we spent a long time searching for a stable algorithm, built an almost complete product around that search, and still failed to prove that the signal engine could consistently improve on the market.
Both versions are true.
The second one is more useful.
This is not a story about how we “found nothing.” It is a story about how a project that could build almost anything finally learned what it should believe.
It began in Excel
The original hypothesis was simple enough to explain without a diagram:
If bookmaker prices contain a collective estimate of team strength, expectations, and new information, then similar market states in the past may help us understand future events.
The first laboratory was a spreadsheet.
A few odds columns. Filters. Historical results. And a person willing to stare at rows for several hours while thinking, “All right, what happens if I keep this, add that, and remove the suspicious-looking thing over there?”
We selected several values, found earlier matches with similar conditions, and counted which side had won more often. If Home had one more historical win than Away, Home became the predicted side. If Away led the count, Away did.
Mathematically, it was naive.
Scientifically, it was useful.
At that stage we were not trying to build a rocket. We were checking whether the material contained fuel.
Then came the opposite idea
Very quickly, a second hypothesis appeared: ANTI.
If a method is consistently wrong, perhaps its opposite is useful.
The intuition sounds almost perfect. If a compass always points south, turn the map around.
The market adds an unpleasant footnote. A bad normal rule does not automatically become a good opposite rule. The other side has its own price, its own break-even point, and the same bookmaker margin waiting at the door. Normal and anti can both lose. The rakes are symmetrical; only the handles are painted differently.
Still, the idea mattered. It taught us to inspect both a conclusion and its negative reflection.
Sometimes the opposite side was genuinely more interesting.
Sometimes it was just a fresh set of rakes, carefully repainted in a more optimistic color.
When Excel stopped fitting inside Excel
Manual analysis met the obvious enemy: combinatorics.
With dozens or hundreds of columns, even simple three-condition combinations multiply faster than rabbits given administrator access to a server.
A person can check ten variants. With heroic patience, perhaps a hundred. Not millions.
So we wrote scripts.
What would have taken years by hand could now run in a few hours. At one point, the system checked roughly seven million three-condition combinations, assembled historical samples, made majority decisions, and recorded normal and anti results.
Technically, this was a major step forward.
Methodologically, it introduced the first major trap: computational speed began to impersonate question quality.
If you test seven million combinations, some of them will look wonderful. Not necessarily because you found a law of nature. Sometimes because seven million lottery tickets are statistically capable of producing several very persuasive souvenirs.
We built a powerful searchlight and did not immediately notice that we were pointing it in every direction at once.

An ocean of reports
The next step seemed logical.
If there are many algorithms, select the best ones. Then combine several of the best. Then check whether they agree. Then create groups, contours, towers, policies, selection layers, deduplication, train/holdout/live splits, and additional controls for the controls.
The project grew.
New stages appeared. New folders appeared. New names appeared. Some sounded as if an intergalactic committee had already approved the future and was waiting only for procurement to sign the form.
The real problem was not the names. It was that one simple question became increasingly difficult to answer:
Why should this exact result repeat on genuinely new data?
Sometimes five different constructions supported the same conclusion. Later we discovered that they were not five independent opinions. They were one idea wearing five shirts.
That is not consensus. It is a family council attended by identical twins.
We were no longer selecting mechanisms. We were often selecting the prettiest outputs of systems built from the outputs of earlier filters.
The ocean became impressive. The currents remained poorly mapped.
When the product outran the evidence
In parallel, we built a serious public shell:
- automatic signal generation;
- sealed snapshots;
- SHA-256 integrity hashes;
- a timestamp-proof architecture;
- a public Signal Key;
- an archive and result audits;
- Discord delivery;
- server operations and recovery instructions.
Engineering-wise, it was a real rocket.
The order of construction was unusual: we completed much of the hull, dashboard, black box, and emergency manual before we had finished proving the engine.
Or, less heroically: we built an excellent safe before finding the gold.
The safe was not useless. The hash, manifest, archive, and audit work remains valuable. But the sequence was wrong. A private engine should survive strict shadow validation before it becomes a public signal product.
The first public experiment remains visible as a historical artifact. In the archive snapshot marked updated July 1, 2026, it contained:
2 archived keys
16 signals
7 wins / 9 losses
−3.92u flat
−9.58u Kelly QuarterThe sample is far too small to condemn an entire research field. It is large enough to trigger a product smell test: the public layer arrived before the signal engine passed its own examination.
There is another detail worth stating plainly. We designed an OpenTimestamps proof path, and the public explainer describes proofs as part of the snapshot contract. In the July 21, 2026 fact check, all three archived keys showed “timestamp not attached.” So the honest claim is narrower:
We built the commitment, hashing, manifest, archive, and proof architecture; in that fact-checked archive state, the three public keys did not have attached OpenTimestamps proofs.
That discrepancy belongs in the postmortem. A verification layer should not receive magical immunity from verification.
The irony survived intact: our proof-of-publication engineering was more mature than our proof-of-edge research.
Two months as a human cable
AI agents accelerated the work enormously.
For about two months, the operating loop looked like this:
Lead agent formulates a task
→ human transfers it to the coding agent
→ coding agent writes code and a report
→ human carries the result back
→ next stage beginsOn paper, this resembles a compact digital team.
In practice, the human occasionally became a highly responsible USB cable between two very intelligent machines.
The scripts arrived quickly. Tests passed. Reports became detailed. The pipeline advanced.
Understanding did not always advance at the same speed.
That was the warning.
A project is fragile when its owner cannot open one match and explain why it entered a sample, where a number came from, which price was used, and what the latest filter actually tests.
Green tests are useful. A report stamped PASS is useful. Neither replaces human comprehension.
A new rule emerged:
Until we verify it ourselves, we do not believe it.
Not until the test suite is green.
Not until an agent writes a confident conclusion.
Until we can open the rows, reproduce the formula, and explain the mechanism in plain language.
Why we did not close the entire project
At one point, the honest question was whether to stop everything and start something new.
After that much work, asking it was not weakness. It was overdue maintenance on reality.
Instead of discarding the whole project, we changed its shape.
The old system was a monolith: each new thought was welded onto the previous structure. The replacement would be modular:
- opening-price state;
- market movement;
- individual-source behavior;
- global market behavior;
- Market Memory;
- head-to-head history;
- source agreement and disagreement;
- sample support;
- normal and anti interpretation.
Each module could be connected, disconnected, and tested on its own.
LEGO, in principle.
In practice, the instructions were missing, there were 1,275 pieces, and the cat had already carried one important blue brick under the sofa.
The real pivot
The old approach asked:
Which of millions of historical results looks best?
The new laboratory had to ask:
What exact hypothesis are we testing, why should it work, and what result would make us close it?
The difference looks small on paper.
In practice, it separates a search for attractive history from a test of a mechanism.
Before rebuilding, we already had several principles paid for with time, servers, and a respectable collection of rakes.
Search scale is not inference quality
Millions of tests increase the probability of finding a beautiful past. They do not automatically increase the probability of finding a stable law.
A perfect tiny sample is still tiny
Three wins from three matches is not transformed into strong evidence by the elegance of 100%.
A large sample still needs meaning
Five hundred matches can mix seasons, sources, price zones, and market regimes. Volume is protection against some errors, not permission to stop thinking.
Five related columns are not five witnesses
If several metrics come from one historical sample, they are relatives—not an independent expert committee.
Anti is a separate bet
The opposite side has its own price and margin. It is not merely a minus sign in front of normal.
Public honesty does not create predictive quality
A sealed record can prove that a claim existed before the result. It cannot prove the claim was good.
A good negative result is cheaper than a bad positive one
A clean rejection can save months. A false positive asks for another server, another policy layer, and just a little more faith.
White curtains
At one point, we renamed files and changed the project structure. It felt slightly like deciding to transform your life by replacing red curtains with white ones.
Curtains do not repair an engine.
But a new order can make an honest restart possible.
We did not erase the previous work. We stopped treating it as mandatory instructions.
The first rocket never became a finished spacecraft. It did teach us where the hull cracked, where the sensor lied, where a report smiled too convincingly, and why returning to Excel can sometimes be more scientific than launching another server cycle.
Part II begins with the rebuild: one stage, one job; one movement formula; individual and global rulers; explicit equality and missing-data states; a match-centric time contract; and a canonical dataset that could be opened, traced, and challenged without first consulting an archaeological map of Python wrappers.
Globa—our imaginary patron saint of suspiciously beautiful backtests—raised one eyebrow.
AGI was asked to bring the formulas.