Evals + post-training
Evaluation protocols, post-training experiments, and behavioural and mechanistic tests that measure what a model actually does, not just what it scores.
Research lab · Brisbane, AU
Independent research in learning systems.
We explore how AI learns, how to evaluate it, and how to bring useful intelligence into everyday software and devices.
We’re working toward models people can adapt, research they can build on, and AI that runs on their own hardware and inside their own applications. Our aim is to put more of that capability within reach through open research and practical tools.
Evaluation protocols, post-training experiments, and behavioural and mechanistic tests that measure what a model actually does, not just what it scores.
Models that keep adapting after training while retaining what they know, with clear lines between what lives in the weights, what is retrieved, and what must be recalled exactly.
Small systems that learn and act in browsers, on local devices, and inside WASM hosts, with the runtime matched to the task.
Small prototypes that test unfamiliar learning loops, memory designs, agent runtimes, and model structures to find out which ideas are worth building on.
We build and study AI that people can make their own.
Projects in evaluating creative work, model behaviour, persistent learning, and intelligence on local devices.
A benchmark in development for AI-assisted fiction. Given an author’s style and a fixed story outline, can a model set aside its default voice, write like that author, and follow the outline faithfully? Style, outline fidelity, and closeness to the source are treated as separate dimensions rather than folded into a single fluency score.
Developing ways to identify AI models with hidden behaviours that ordinary testing may miss. The aim is to catch sleeper agents before those behaviours emerge in use, and give people better evidence for deciding which models to trust. Early work studies what a model’s internal activations reveal about how it has been altered.
Exploring how AI can keep learning from experience without losing what it already knows. ANNEAL is a lightweight layer that writes new knowledge into a pretrained model, then decides how much of each update to keep. The aim is models that grow with the people using them, light enough to run locally.
Browser-first machine learning built on a Rust/WASM core and WGSL compute. The aim is to train small models on a user’s own GPU without a dedicated training server, making hands-on machine learning more accessible and letting applications learn on device. So far it has been validated on a single browser and device.
A compact ECS-flavoured runtime for agents in applications, games, and simulations, targeting WASM and native hosts. Agents can be classifiers, state machines, or small MLP policies as well as language models, bringing adaptive behaviour to more kinds of software and hardware.
Research collaborations, applied prototypes, evaluation work, or an early idea you want to test: we’d like to hear from you.
Direct contact
shannon@artivus.com.auTell us what you’re working on and where you’d like help.
Brisbane, Australia