Research lab · Brisbane, AU

ARTIVUS

Independent research in learning systems.

We explore how AI learns, how to evaluate it, and how to bring useful intelligence into everyday software and devices.

Harmonograph phase-space traceA decaying multi-frequency plot used as a visual instrument for interacting learning dynamics. f₁ 2.01 · f₂ 3.00λ 0.0035SIG/01

Focus areas

We’re working toward models people can adapt, research they can build on, and AI that runs on their own hardware and inside their own applications. Our aim is to put more of that capability within reach through open research and practical tools.

01

Evals + post-training

Evaluation protocols, post-training experiments, and behavioural and mechanistic tests that measure what a model actually does, not just what it scores.

02

Continual learning

Models that keep adapting after training while retaining what they know, with clear lines between what lives in the weights, what is retrieved, and what must be recalled exactly.

03

Lightweight multiplatform intelligence

Small systems that learn and act in browsers, on local devices, and inside WASM hosts, with the runtime matched to the task.

04

Architecture experiments

Small prototypes that test unfamiliar learning loops, memory designs, agent runtimes, and model structures to find out which ideas are worth building on.

Our approach

We build and study AI that people can make their own.

Our research explores how AI can keep learning, support creative work, and run on everyday hardware. We’re interested in giving people more control over the systems they use, and better ways to judge what those systems can actually do. We plan to release tools so others can explore these questions with us.
Evaluation loop instrumentHYPOTHESISVERDICTCONTROL / SIGNAL
  • Learning beyond initial trainingNew knowledge and changing needs shouldn’t always mean starting over. We’re exploring how models can adapt over time while retaining useful capabilities.
  • Evaluation for real workChoosing and improving a model takes more than a leaderboard score. We build evaluations around the qualities that matter in use, including creative work that conventional benchmarks struggle to capture.
  • Intelligence on your hardwareLightweight tools for running learning systems in browsers, on personal devices, and inside applications, so developers control where and how intelligence runs.

Research index

Projects in evaluating creative work, model behaviour, persistent learning, and intelligence on local devices.

IN DEVELOPMENT

AuthorBench

A benchmark in development for AI-assisted fiction. Given an author’s style and a fixed story outline, can a model set aside its default voice, write like that author, and follow the outline faithfully? Style, outline fidelity, and closeness to the source are treated as separate dimensions rather than folded into a single fluency score.

Style · Outline fidelity · Creative work
IN DEVELOPMENT

Sleeper-Lens

Developing ways to identify AI models with hidden behaviours that ordinary testing may miss. The aim is to catch sleeper agents before those behaviours emerge in use, and give people better evidence for deciding which models to trust. Early work studies what a model’s internal activations reveal about how it has been altered.

Backdoors · Activations · Model trust
IN DEVELOPMENT

ANNEAL

Exploring how AI can keep learning from experience without losing what it already knows. ANNEAL is a lightweight layer that writes new knowledge into a pretrained model, then decides how much of each update to keep. The aim is models that grow with the people using them, light enough to run locally.

Continual learning · Fast weights · Local
EXPERIMENTAL

WebGPU ML

Browser-first machine learning built on a Rust/WASM core and WGSL compute. The aim is to train small models on a user’s own GPU without a dedicated training server, making hands-on machine learning more accessible and letting applications learn on device. So far it has been validated on a single browser and device.

WebGPU · Rust/WASM · On-device
ACTIVE PROTOTYPE

Snapdragon ECS

A compact ECS-flavoured runtime for agents in applications, games, and simulations, targeting WASM and native hosts. Agents can be classifiers, state machines, or small MLP policies as well as language models, bringing adaptive behaviour to more kinds of software and hardware.

ECS · WASM · Agents

Start a conversation

Research collaborations, applied prototypes, evaluation work, or an early idea you want to test: we’d like to hear from you.

Direct contact

shannon@artivus.com.au

Tell us what you’re working on and where you’d like help.

Brisbane, Australia