About
About
For most of my Year 12 peers, February 2026 marked the beginning of their final and most intense academic chapter. For me, it marked the exact time I decided to throw out the conventional Year 12 script and write my own.
Being in the unique and fortunate position of completing enough final exams in Year 11 to secure my tertiary education, I was faced with a hard choice of what to do in 2026. I could have skipped Year 12 and gone straight to university one year earlier than the rest of my peers. I could have decided to complete Year 12 to the best of my academic abilities so I could obtain higher academic credibility and graduate with my cohort. Instead I spent the year working on artificial intelligence.
Before 2026, I always thought that I would end up working in the financial sector. My dream was to work as a portfolio manager, helping manage money for my clients and working with the financial markets, which was my passion for a very long time. I am not ruling out a return to it, and the financial literacy I built while exploring the industry is one of the most useful things I have.
The year started commercially. I built QuoteReady, an offline-capable quoting tool for electricians and plumbers, and sold it pre-revenue in April. I built a database of more than 1700 modern methods of construction companies across Australia and New Zealand. Then I spent months trying to automate construction takeoff with vision-language models, and failed to reach the accuracy a real business would need. Writing up why I failed taught me more than the three successes had. The ceiling was not the model. The information that determined the answer was never in the drawings to begin with, and no amount of capability was going to recover it.
That write-up changed what I work on. The question I had accidentally spent months on was not how to build the product, but how to tell whether a system was doing what everyone assumed it was doing. Since then I have built DocQAGroundingEval, an open-source harness measuring whether frontier models actually rely on the document in front of them or quietly fall back on what they already believe, and I spent a 48-hour Apart Research sprint probing whether a hidden loyalty planted in a model can be detected from the outside. That one placed in the top 25% of 179 projects, and the most useful thing to come out of it was a control task proving my own headline result meant nothing. I am now working through the ARENA curriculum from the beginning to close the gaps between what I can direct and what I can build unaided.
The through-line, as far as I can tell, is that I keep building the test that could prove me wrong and then reporting what it says. A positive result is only as good as the negative case you tested it against. That habit is worth more to me than any single finding I have published.
Where this ends up is still open. I would like to build something that matters at scale, and I am spending part of this year with EnergyLab's Ignite program looking for a problem worth building a company around. I would rather leave that program with a real problem and no company than with a company and no problem. What I am more certain of is the shorter horizon: I want to work on AI safety and evaluations, with people who will tell me when I am wrong faster than I can work it out alone.