Work with the lab

Exploring new ways to build, understand, and interact with artificial intelligence.

Apperson Labs is a small, independent research lab. We publish models, open-source tools, benchmarks and papers on how people and intelligent systems work together.

ResearchModel

Arbiter

A family of small decision models. Given evidence and a short list of options, Arbiter says which option holds and how sure it is.

Read the model release

Latest from the lab

At Apperson Labs, we build to understand how people and intelligent systems work together.

How Arbiter compares with models many times its size

  • Arbiter M8B, choice readout
    93.1%91.9 to 94.2
  • Arbiter S2B, choice readout
    89.4%88.0 to 90.8
  • Frontier baselinelarge hosted model, prose answer
    94.0%92.9 to 95.0
  • Open 8B baselineuntuned, prose answer
    81.2%79.4 to 83.0
  • Open 2B baselineuntuned, prose answer
    71.6%69.5 to 73.7
Accuracy across 1,840 decisions, with 95% intervals.

Illustrative data. These figures show the shape of the evaluation and will be replaced by measured results at release.

Loop or progress

A retry that is really a poll

gh run watch 8841 exits 1 with “run still in progress” for the fourth time in six minutes. No files changed between attempts.

  • Polling an external jobcorrect answer0.91
  • Stuck on the same failure0.07
  • Cannot tell0.02

A retry counter would have stopped this at three. The job finished on the sixth check.

Apperson Labs / Noetic

Noetic

An open-source TypeScript agent framework that breaks agent patterns into eight composable step primitives.

Everything in Noetic is a typed, serializable step. There are no base classes to inherit from and no hidden control flow. A loop can contain a branch, which can contain forked, spawned agents.

loopuntil done
step.llmplan
branchneeds research?
forkall, in parallel
spawnisolated context
step.tool
step.llm
spawnisolated context
step.claudeCode
step.runmerge
A research agent written with all eight primitives. Each box is a typed step.

Research becomes experiments, experiments become releases

Nothing here starts as a product. Follow any piece of work back to the question it came from.

ResearchExperimentModelBenchmarkReleaseJudged loopsChoice readoutsNoetic Rules judgeStardrive crewArbiter SArbiter MDecisionsNoetic Rules betaResearchExperimentModelBenchmarkReleaseJudged loopsChoice readoutsNoetic Rules judgeStardrive crewArbiter SArbiter MDecisionsNoetic Rules beta
  • Judged loops leads to Noetic Rules judge
  • Choice readouts leads to Noetic Rules judge
  • Noetic Rules judge leads to Arbiter S
  • Noetic Rules judge leads to Arbiter M
  • Stardrive crew leads to Arbiter M
  • Arbiter S leads to Decisions
  • Arbiter M leads to Decisions
  • Decisions leads to Noetic Rules beta

Stardrive

A game whose crew are long-running agents with memory, temperament and their own opinions about your orders.

Why a lab makes a game: nothing else asks as much of an agent for as long.

  • Agents that run for the length of a save file
  • Memory that changes how a character behaves
  • Believable responses inside a frame budget
  • Inference cost a player never has to think about

Working on a similar problem? Work with the lab.