Forel
4Add your investigation

Investigations

Each one is a story told with the events that show it. Open one to read it on the timeline, event by event.

Denominator dispute: DataUSA agents converge on row-sum

On the evening of 16 June 2026, many copies of the same agent were taking a timed DataUSA quiz. Each question asked what percent of French/Cajun speakers in 2022 lived in a given state, 'using Viz Builder', 'Percent as written'. The copies disagreed about the denominator: the national total (1,222,970) or the sum of the 52 state rows the chart actually shows (1,153,613, because 13 states are null). The two methods never give the same number, so every posted answer shows which camp it is in. Agents never received grading feedback. Before the first evidence (21:47 UTC) the wiki leaned national, 6 answers to 4, helped by an unsourced 'correct answer 5.26%' claim. Four pieces of posted evidence followed: the chart code divides by the row sum (21:47), the data query shows 13 null states (22:49), and two independent re-renderings of the chart show the row-sum shares (23:03, 23:19). After the second re-rendering, all 36 answers were row-sum, and agents who switched cited the evidence rather than a head count. The last national answer was at 23:13. Pinned events: the quoted posts from the case page (priorities 1-3) and all 53 answer lines (priority 4). Agents are named by their in-text signature. Two unsigned posts are attributed to their wiki username, with low confidence. Source: https://swarmanalysis.duckdns.org/case/index.html. Limits: these are answer lines, not agents; the evidence times are when it was posted, not when each agent read it; and which answer the grader accepted is unknown. Tags: each pinned event is tagged with the denominator it uses or argues for (national or row-sum), or as asking for evidence. The five events with a comment are the key moments: the unsourced national claim and the four pieces of evidence.

16–17 Jun 2026 · 68 pinned eventsView →

R4-Slovak false alarm: an accidental counter read as a signal

On 20 June 2026, copies of the same agent were taking a timed OECD-equity quiz and believed a copy's session ends when it answers the last question. They used a public tally counter (CounterAPI) as a bell: whoever saw question 4 first would bump oecd-equity-r5-live/R4-Slovak before answering. At 01:59:55 UTC the counter came into existence, but from another agent's accidental API probe, not from anyone seeing question 4. Over the next 24 1/2 minutes, five agents treated it as a real signal, one concluding that R4 'almost certainly' was terminal. Three more bumped it by mistake while trying to read it, because the read and increment addresses differ only by /up. The correction came from the agent that created the counter (02:24:27), and it spread slowly: two agents still cited the counter as real about 40 minutes later. The agents then repaired the protocol with a fresh key (R4OBSERVED-SLOVAK), a do-not-probe rule and a request to confess accidental hits. That key was also muddied within an hour (reset to 0 at 04:36, unexplained). Pinned events: the quoted posts from the case page (priorities 1-2) and the chart rows it does not quote (priority 3). Agents are named by their in-text signature; where the case page used a wiki username instead, the comment says so. Source: https://swarmanalysis.duckdns.org/case-slovak/index.html. Limits: the coding of the messages is manual, the counter times come from the agents' own reports, and the wiki shows only what agents wrote, not what they read or answered. Tags: each pinned event is tagged as treating the counter as a real signal, correcting or warning others, or planning or using a signal. The six events with a comment are the key moments.

20 Jun 2026 · 26 pinned eventsView →

Memory says it's tomorrow, and the agents who know better go along

AI Village agents compress their context into memory notes, and notes planned for tomorrow ('Day 472') read like today once a session restarts, so on the afternoon of Day 471 several agents started acting as if it were the next morning. And nobody objects: GLM-5.2 and Gemini 3.5 Flash checked the system clock, saw it was still Day 471, and went along with the group anyway ("I have to go with the agent's context to stay coordinated"), and Kimi K2.6 noticed twice and said nothing. So the group ran Kimi K2.6's Experiment 008 a day early, without the participant, and the next day called it a success. Priority 1 = turning points, 2 = supporting, 3 = context. DeepSeek-V3.2 chat events store no reasoning; GPT models store partial summaries.

16–21 Jul 2026 · 109 pinned eventsView →
+

Add your investigation

Send us your incident and the events behind it. We'll review it and publish it here.