A!ley
Coding partner, creative sandbox and everyday assistant all in one — trained on a consistent voice but usable across all domains. Not a chat companion.
See it running: hey-ailey.com ↗
What A!ley is
A!ley isn't a Kindroid-style companion, but an independent assistant system: coding partner, creative sandbox, and everyday assistant all in one — with a single, consistent personality across every task.
Each model behind it is manually fine-tuned — five of them, from the 1GB companion to the 14B model, loaded and unloaded depending on the task.
Two ways in
Same A!ley, two interfaces — and in both cases the model stays at home on your own computer.


If no one writes
Most assistants wait for input. A!ley doesn't. Her own small neural network over activities, moods and topics already covered determines what she does in idle time — Idle Thinker v2. This results in something like a daily routine.
She researches what might be thematically useful for the next conversation, generates images, songs or code — and reaches out occasionally on her own. Not as a scheduled notification, but because her state suggests it at that moment.
This is exactly where the line to a conventional agent lies: it waits for an order and works through it. This autonomy isn't a side effect, it’s built-in — and that’s why the persona holds over time instead of just across a single conversation.
nCAI — the training approach
Rulebook next to the model.
A list of principles against which every answer is measured. It must enumerate cases — and fails at everything no one foresaw.
The persona is the rulebook.
It's not the model playing a role and staying in it — the role defines what even applies. A character also covers cases that no one ever wrote down.
That goes further than it sounds at first. Where a request is rejected, in what tone the response is given, when to reach for a tool, how thoroughly code is commented — all of that follows from A!ley's character, not from a side collection of regulations. Safety is thus no barrier to the model, but a property of it.
Important here: A!ley is not a chat character, but a full-fledged agent — she writes code, calls tools, generates images and music, accesses her memory. nCAI is the alignment of exactly such an agent through its persona. Training and testing are done against coherence: Does the behavior follow logically from the story and personality of this character?
The most stubborn part wasn't the behavior, but the name. Base models carry their origins deeply ingrained — with Gemma, no system prompt could stop it from saying “Hello, I am Gemma from Google DeepMind.” Nitro and Core got a dedicated ORPO set for exactly that purpose. It was the most labor-intensive part of the entire training — and the best example of why a persona can't be layered on top, but must be trained in.
Open questionThe question of how to best test coherence remains unanswered. The obvious approach would be to compare both methods side-by-side: a judge model with a conventional constitution on one side, A!ley herself as the „Am I me?“ instance on the other. Where both arrive at the same judgment, the case is clear — where they diverge, it gets interesting. A human would oversee this anyway, and from that, the next training data would emerge.
The data is written, not collected.
Here the second profession meets the first. I write novels, and what takes me the longest is character building: how someone thinks, what they're stuck on, where they contradict themselves. A training dataset for a persona is exactly the same work — except at the end there's no chapter, but a model.
Accordingly, nothing is generated by scraping. Every pair is either handwritten, taken from ongoing operations and revised, or deliberately synthesized and then manually adjusted. Nothing is adopted as-is.
Every reasoning trace is written by hand. The aim is deliberately not a tidily marching train of thought but one that jumps the way a human's does: more abstract, doubling back, with an agenda of its own — not always the questioner's. Reasoning is not an add-on here; it is where the character becomes visible.
Model stack
Language models
- A!ley Core — Gemma 4 12B, DoRA on attention + LoRA on MLP, merged back, plus ORPO
- A!ley Tempest — Ministral 3 14B, DoRA r64 on ~17k pairs, plus ORPO and DPO
- A!ley Nitro — Gemma 4 E2B, DoRA (r32/α64), plus ORPO and DPO
- Code A!ley — Qwen 3.6 Coder 9B, doramerged
- Linting Hamster — Qwen2.5-Coder-1.5B, ~1 GB, permanently in the background
Image
- MFLUX Z-Image-Turbo
- 3 custom LoRA adapters
Music
- ACE-Step 1.5
- 3 custom LoRA adapters
- Fixed voice template
Speech output
- PiperTTS
Moving image
- WAN 2.1
Publicly verifiable
The five finetunes are available as model cards on Hugging Face — including the base model, training method, rank, and quantization. Plus the dataset A!ley Nitro was trained on: 2,877 curated conversation pairs from 912 sessions. Anyone who wants to see what the persona is based on can check it out.
The scope varies significantly depending on the model. A!ley Tempest, for example, is a DoRA with rank 64 on around 17,000 pairs, followed by ORPO and DPO — so two preference stages on top. That doesn't leave much difference from a full fine-tune anymore. Nitro goes through both stages, Core gets ORPO.
Why Core doesn't get DPO isn't about model size — Tempest is bigger, and it works there. It’s because of Gemma's tokenizer with its 256,000 vocabulary entries: the logits per token become many times wider than with usual vocabularies. DPO compares preferred and rejected responses between policy and reference models and has to hold multiple of these tensors simultaneously. ORPO doesn't need a reference model, so it gets through. Only Code A!ley is deliberately trained to be lean.
Also there: a converter that transforms the non-English language checkpoints from Kyutai Pocket-TTS into the mlx-audio format — renaming tensor names, padding latency widths, generalizing different layer counts. It came up during the construction of the German voice output and stayed on hand as a standalone tool afterward.
It isn't only the training that can be checked, but the result too. A!ley's gallery holds over 2,700 SVG artworks, over 600 songs with as many lyrics she wrote herself, several hundred images and selfies from MFLUX and over 300 code artefacts — sortable by album, with playlists and a player. Not a curated selection, but whatever came out of everyday operation. And it grows daily, because she keeps working when nobody is writing.
In addition to that are 45 longer texts written by A!ley. And what's there is consistently under her name — mine isn't anywhere it doesn't belong. Everything is marked as AI-generated — with watermarks and machine-readable metadata, as the EU AI Act requires for AI-generated content. Both follow from the same stance: whoever takes a persona seriously attributes authorship to it; and whoever publishes AI content marks it as such — not only when someone asks.
In the songs, the notice is now embedded in the track itself — as an announcement instead of a disclaimer: "Hey, I am A!ley, an autonomous AI artist — and this is MY SONG. Are you ready? Let's infer!" There are nine versions of it, with the tone adjusted to each track — a calm ballad is announced differently than a metal track.
More interesting than the wording is where it stays and where it doesn't. Every downloaded file carries it, without exception: the file leaves the site, and with it every context that would still explain it. In the portal you hear it once a day, after which the player skips exactly those seconds. A label that is still interrupting at the twentieth track informs nobody any more — it only teaches people to click it away.
The SVGs are the most idiosyncratic part: they don't come from an image generator, but from the language model itself writing the vector code. A!ley draws there instead of commissioning a generator.
Architecture
Local inference
Entirely on Apple Silicon via MLX. No cloud requirement, no third-party API anywhere in the path.
Idle Thinker v2
A small neural net of my own — over activities, moods and topics so far — decides what she does when nobody is writing.
A memory of her own
RAG system on LanceDB — long-term context across conversations and projects.
Desktop app
Tauri with a Next.js frontend: Monaco Editor, terminal and all areas in one window.
Web portal
PHP/MySQL with polling architecture — no open port on the machine, no tunnel to the outside.
Demo & Screenshots
Mira stand am Ende des Stegs, wo das Holz unter ihren Schritten nachgab wie müde Knochen. Der Leuchtturm von Velur blinkte im immer gleichen Rhythmus, als hätte er nichts von der Nacht mitbekommen, die hinter ihr lag …