| Felix Halim .NET |
Friday, October 9, 2026
Monday, August 3, 2026
Solving 3,000+ onlinejudge.org problems in 3 weeks using modern LLM coding agent
The onlinejudge.org is one of the oldest online judges (see Grokipeida for the history).
My first submission was 26 years ago (2000-01-14 09:47) and the first 14 years were really formative years for me, using the online judge to develop my critical thinking, algorithm and data structures, as well as coding skills and speed. I've solved 1,521 problems by myself (with hints from the web), but since then I've been struggling to find time to continue this developments.
In the recent advent of LLM (large language models), I'm curious how far can a modern LLM coding agent solve problems from this online judge? so I did an experiment to use Antigravity ($20 subscription) and Claude Code ($100 subscription) and ask it to solve as many problems it can. In summary, Antigravity (Gemini 3.6 Flash) and Claude Code (Opus 5 + Fable 5) together can solve additional 3,354 of my unsolved problems in 3 weeks. FYI, it did write most of the solutions by itself before falling back to the web for hints when stuck.
So, it's true that modern LLM coding agent can solve about 97% competitive programming problems (at least those problems that already circulated in the web) easily and really fast too. I would (approximately) say that with Antigravity + Gemini 3.6 Flash you can use it to solve 90% easiest problems, the next 6% harder problems with Claude Code + Opus 5 and the last 1% hardest problem with Claude Code + Fable 5.
Solving 1 problem every 9 minutes
You may be wondering how did I use the LLM coding agent to solve 3,354 problems in 3 weeks (i.e., 1 problem every 9 minutes)? Did I manually copy paste each problem to the Claude Code every 9 minutes for the entire 3 weeks? Yeah... no. I used skills!
Skills for solving a problem
A skill is just a markdown file containing a tutorial of how-to do a certain task for the LLM coding agent to follow. You just write the skill once and then you can refer to it again and again in a prompt. For example, when submitting a C++ solution to an online judge, you need to first login to the onlinejudge.org, then open the quick-submit page then fill in the forms specifying the problem number, language, and upload the C++ solution. To create a skill, you can just prompt Claude Code to write a skill to do that and save it in a skills folder named "skills/submit-problem". Next time, you can just prompt submit problem 12345 using the skill without having to re-explain the steps. In the same way, you need to create skills to lookup the submission status of the solution to the problem that you've just submitted. You can also create a skill to solve a problem: first download the pdf of the problem, parse the input and output, save it to {number}.in and {number}.out, write C++ solution and save it to {number}.cc and then write down the editorial on how to solve the problem and save it to {number}.md. Yes! you can ask LLM agent to explain the solution too!
Solving the problems in parallel
Once you have the solve, submit, and check submission status skills you can then prompt: "please solve and submit problems numbers A, B, C, ... in parallel using 10 subagents until they got AC." Be careful to issue this prompt, it will drain your LLM quota quickly! only do this when you are sure the skills are well tested first! Also you will learn that doing this is not efficient since your subagents tokens will be wasted if you are out of quota in the 5-hours window. It's better to never run of quota in the 5 hours window so your subagents never need to be restarted.
Which LLM to use?
Now, when you want to run the LLM coding agent for long running tasks over 3 weeks, you want it to continuously running while you are sleeping! From my experience, Gemini 3.6 flash is not good for long running tasks: it stops in about 1 hour and need to be kicked-in-the-butt to "continue the work". There's a workaround to ask it to run cronjob to self-kick every 5 mins, but you'll discover that Gemini will often forgot that it run expensive processes and didn't terminate it before moving on to the next problem, this may crash your working computer! You can add to the prompt for Gemini to proactively cleanup the zombie processes, but it still often forgot after a while!
Claude Code with Opus 5 is far better for the long running tasks! It can run overnight if you tune the token consumption so that you are not to out of quota for the 5-hours windows. Claude Code with Fable 5 is just too slow for solving problems and burning too much tokens, this is better reserved to solve problems that are unsolvable by Opus 5.
Competitive Programming Future
With LLM coding agents that powerful, what will happen to the future of competitive programming? Well, I think it will be even more competitive! The situation is similar to Chess or Go where there exists more powerful engine than human, but the competition between human still going strong. Human can learn even faster with the help of the engine. The engine will become cheaper and better in explaining the solutions (as long as the human still have willingness to learn).
This also means that cheating in online contests is now getting easier than ever! Contest organizers will need to figure out ways to make the competition fair for the players.
Looking at this World Ranklist below, the top ranked user is still human (Josh Bao). FYI, I did use my own user name (felix_halim) to run the experiment, and I've ended my experiment (since Claude Fable 5 also gave up on the remaining unsolved problems saying that they are research problems and too difficult to solve). So at this point, maybe human is still winning :)
![]() |
Sunday, December 14, 2025
The Autonomous Horizon
Note: this is an infographic produced using Gemini 3 pro deep research with the following prompt: "explore the future of self driving car. will tesla be the only winner? will other car manufacturer license tesla fsd? how hard is it to license fsd? what are the requirements? will self driving be like LLM where other companies can catch up quickly and becomes commoditized? if there are multiple winners who are the candidates? will there be like "android of fsd" where fsd get open sourced? if car manufacturer don't want to license fsd from tesla, what alternative do they have? what is the likely hood of other car manufacturer to build their own fsd technology? can they distill the fsd model from tesla? or they have to do the hard work of collecting real world data for fsd or can they just use simulation?"
THE AUTONOMOUS HORIZON
Will Tesla take it all? Exploring the data moats, licensing barriers, and the future battleground of Full Self-Driving.
Data is the New Oil, and Tesla Has the Pipeline
To solve self-driving, you need to capture the "long tail" of weird edge cases (e.g., a person in a chicken suit crossing a highway). Simulation can only guess what it hasn't seen.
Tesla's fleet of millions of consumer vehicles provides a data feedback loop that traditional robotaxi fleets (like Waymo) struggle to match in pure volume. This "Data Moat" is the primary argument for Tesla's potential winner-take-most outcome.
Cumulative Autonomous Miles (Est.)
*Logarithmic scale visualization for impact comparison
Why Other Brands Can't Just "Install" FSD
Licensing Tesla FSD isn't like installing Android on a Samsung phone. It requires a complete architectural overhaul. The hardware and software are tightly coupled.
Hardware Mismatch
Most cars use supplier-grade cameras and radars. FSD is trained on specific Tesla vision inputs. To license FSD, Ford or GM essentially has to build a Tesla clone.
The Black Box Problem
Automakers want to "own" the experience. FSD is an end-to-end neural net. You can't tweak it easily to drive "more like a BMW." It drives like a Tesla.
Validation Costs
Validating the software on a new chassis takes months or years. It's not a simple plug-and-play API integration.
Is There an "Android" of Self-Driving?
While Tesla pursues a vertical Apple-like strategy, others are vying for the platform role. Who has the best shot at being the alternative?
Competitor Capability Matrix
Waymo (Google)
**Strategy:** Geo-fenced Robotaxis using Lidar + Maps.
**Pros:** Extremely safe, proven driverless operation today.
**Cons:** Doesn't scale easily to consumer cars or random locations.
NVIDIA + Mobileye
**Strategy:** The Arms Dealers. Selling chips and vision stacks to everyone else.
**Pros:** The "Android" path. Low risk for OEMs.
**Cons:** Fragmentation. Data collection is slower than a unified fleet.
Chinese EVs (XPeng/Huawei)
**Strategy:** Fast follow + aggressive domestic mapping.
**Pros:** Innovation speed is matching Tesla.
**Cons:** Geopolitical barriers to Western markets.
Comma.ai
**Strategy:** Open Source / Consumer Hardware Retrofit.
**Pros:** Cheap, runs on many cars.
**Cons:** Limited authority over car controls; niche market.
Will FSD Be Commoditized Like LLMs?
With Large Language Models (LLMs), we saw rapid commoditization (OpenAI -> Llama -> Mistral) because text data is available on the open internet.
**Self-driving is different.** You cannot scrape the internet for physical driving intuition. You need video stamped with steering angles and acceleration data. This creates a much deeper moat.
- LLM Barrier: Compute Cost (Medium), Data Access (Low)
- FSD Barrier: Real World Data (Extremely High), Regulation (High)
Training Data Value Composition
Legacy Auto's Hard Choice
Distill: Can't copy Tesla directly.
Risk: Bankruptcy if failed.
Requires hardware redesign. Swallow pride. Pay royalty.
Become a commodity hardware assembler. Loss of differentiation.
