AgentOB · Build log 01
My first thought about an AI game observer was straightforward: I believed AI was good at prediction. If we continuously supplied enough real-time context, could it anticipate what would happen next?
Applied to a game, the question became: could it anticipate where the next worthwhile shot would be, decide whose perspective to show, and do the work of an observer?
At that point, this was a hypothesis to test. How far AI's ability to predict text could carry into engagements and camera choices needed actual implementation. I was also thinking about star players, their background, and the story of a competition. A shot might deserve time because of a possible comeback, or because viewers were already following that player's situation.
Then Jev, a decision model, was released. Its ability to return decisions directly to software prompted us to think further about the observing idea. Since the match takes place inside a digital system, could we give AI the data directly and let it choose from player positions, actions, and events?
This brought a distinct realization for me. When thinking about how to solve a problem, I tended to follow human habits: we look at an image, understand what is happening, and explain our judgment in language. Designing a task for AI could follow the same habit, asking it to look at the screen first and then explain what it wanted to do.
A machine can use a different information path. Game states already exist as data. It can read those states directly, process numerical and logical relationships, and return a choice that software can execute. The image a person sees on screen and the input a machine needs for a decision can be designed separately.
That gave the original idea a concrete entry point. We began specifying which information the program could read and how it would choose a camera, then played those choices back to see whether viewers could follow the match.
AgentOB entered engineering through that opening. It now reads CS2 match demos, replay files that preserve the course of a game, selects players, switches cameras, and plays continuously in a 2D view. We can also compare versions on the same passage of a match.
Once the program ran, new problems began to appear. It switched the camera to a player, then cut away half a second later. It selected someone. I barely had time to see what was happening there. Before adjusting the scores for selecting players, I had to explain what this shot should let a viewer see.

First, constrain a program that knows the ending
A match demo contains a game that has already happened. A program can easily gain access to later outcomes.
Before implementation, I set a requirement: each decision could use only what was known at that moment. I wanted to see whether it could anticipate a good shot, so later events had to be kept out.
That requirement became limits on inputs and time. It also shaped how we read results. When the program selected someone early, we had to check which facts supported that choice. Knowing who would get a kill and then selecting them would leave the original question unanswered.
Later, we separately studied a condition that allowed five seconds of upcoming events. It lets us explore observing with a delay, and must remain separate from versions using only current information. The five-second window works on an existing replay. Live broadcast buffering has not been implemented.
This became a working habit for the project: specify what the program may know and what it should do, then inspect what it actually did. Even an appealing camera sequence needs an account of how it was chosen.
I brought the program specific things that felt wrong
As the project developed, I began to experience more concretely how much work it takes to put judgments people make through experience into code.
For example, we used distance between players as one signal for a possible engagement. I brought up the middle doors on Dust2. Two players can be far apart, yet a sniper and someone crossing the doors can form a confrontation worth showing early. Distance leaves out the position, the weapon, and the route being taken.
That example entered the discussion as a counterexample. Turning it into a reliable rule still needed validation. It helped me point to something missing from the program's judgment, while exposing what I had yet to explain precisely myself.
I have been building the project with several AI collaborators. I set directions, bring in passages and counterexamples, ask them to organize proposals, write code, and run comparisons, then bring the results back for another look. I also decide what deserves more work and what should pause.
Giving a judgment about a shot is often quick. Explaining which facts it depends on and when it fails takes longer. AI can keep writing the implementation; I have to keep supplying the grounds for the judgment.
The recent experiments gradually took shape around a few concrete passages.
What can you see in half a second?
One of the first problems was fleeting shots.
The camera switches to a player, then leaves half a second later. The system may have selected someone involved in an important event, but the viewer has barely had time to place them. Several such choices in a row fragment the sequence.
My requirement for this passage was clear: give people enough time to understand what they are seeing. We added a minimum hold and started by testing two seconds.
Restricting short shots exposed another cost: while the camera stays here, something may already be happening elsewhere. Counting switches can make the new version look calmer. Watching the match also means checking what it missed.
A request as simple as hold the shot longer acquires a cost in code. Two seconds is a parameter we have tried. The right duration still depends on the actual scene.
A coordinated play needs two players
In another passage, Wicadia throws a flashbang and XANTARES then takes the fight.
If the camera only picks up XANTARES during the engagement, viewers see the execution and may miss the preparation supplied by his teammate. Following the thrower first, then switching to the player taking the fight, can connect the two parts.
We have a working expression of that idea in one case. In an offline experiment allowed to read five seconds of upcoming events, the camera follows Wicadia, switches to XANTARES after the throw, and stays for part of the engagement.
The process also exposed an incorrect event association. We checked the raw events and corrected the condition. Connecting the two players requires more specific evidence than an impression that this was teamwork.
That camera path now connects. Its effect on the whole match still needs checking. Giving it time can cost another sequence its shot. The compared strategies differ in several ways; this local case does not establish better overall viewing quality.
Showing the two related players in sequence gives viewers a chance to see the coordination. Doing this reliably in other matches remains unfinished.
A file on waiting: 38.5 seconds, unresolved
A roughly 38.5-second hold on Spinx remains unresolved.
The easiest response is to cut away when a timer expires. But the absence of shooting does not establish that holding an angle has no value. The shot may be waiting for someone to take a route, or sustaining an expectation that has yet to pay off. Other players are moving, firing, and using utility during that time.
Chasing every new activity can bring back the fragmented cuts.
We tried giving the camera a task: what process it is following, when that process finishes, and what change would justify handing attention to something else. The long hold persisted. More lookahead did not fix it either.
The failure is worth keeping. It makes the questions concrete. Under what conditions does the wait retain value? Does the viewer know what it is waiting for? What change would remove the reason to continue?
The next revision has to answer those questions. Another timer still leaves those 38.5 seconds unexplained.
Connecting a model leaves its benefit to be tested
Jev did enter the project. In the early setup, rules picked an event and the model chose a perspective from the associated players. Later, we added camera history and let it help choose which process to follow.
That round of local comparisons still showed no added benefit, so we paused this line of work. The default observer continues to use rules. The model code and results are preserved.
This round tells us that the current integration has yet to deliver the improvement we want. Before reopening it, we need an answerable question for the model: which facts it can see, which decision it can change, and how we will judge whether the resulting sequence is better.
Jev initially helped me find an engineering entry point. At this stage, we still need to break the observer's judgments down further.
This exhibit stays on the workbench
AgentOB is still in development and is not open source yet. What exists is an offline foundation for executing camera choices, playing them back, and comparing versions, along with a few pieces of expertise translated into working conditions. Real CS2 camera execution remains unvalidated. Helping viewers consistently understand a match remains unfinished too.
I want to keep recording its development here at the Institute. A passage that is hard to follow gets identified, translated into conditions, and revised. We then check what the revision recovered and what it gave up. Successful cases and unresolved waits stay in the record together.
The record also keeps track of how my questions change. I began by asking whether AI could anticipate a good shot. Later, I started asking what setup a shot needs and when a wait should end. Whenever the program makes a choice, we can return to the same passage and inspect it.
The next step is waiting. I will start with contrasting passages where a shot should continue or end, spell out the evidence each judgment depends on, make a small change, and check it against the same match.
If you watch CS2, or work in observing or editing, I would welcome a specific passage where you felt the camera should cut or stay. Include what you wanted the viewer to see. We can make it the next case to unpack.