Machine learning research & engineering

Lorenzo Iotti

I care about what AI feels like to use, and I build models until it feels right.

I start from an experience I want to exist, and work backwards through the data, the training and the model's internals until it does.

Now 2026

Full-duplex speech on Gemma 4

One model that listens and talks at the same time, deciding every 80 ms whether to stay quiet or speak.

YOU MODEL mm-hmbarge-inyeahmm-hmbarge-inyeah NOW
Gemma 4 12B · Mimi codec

01 — Work

Selected work

  1. 01 2026 Real-timeSpeech

    Full-duplex speech on Gemma 4

    A model that listens and speaks at the same time, so a conversation with it can overlap, pause and change direction the way conversations between people do. It's built on Gemma 4 12B, which had never produced a second of audio: every 80 ms it hears raw microphone audio and decides whether to stay silent or speak, with no ASR or TTS in the voice loop. It learned to speak from under 100 hours of synthetic conversation, a tiny fraction of what speech models are usually trained on, and it runs in real time on a MacBook. It delegates work to external agents while staying in the conversation, and a separately trained continuous acoustic stage gives it different voices without retraining the conversational model.

    Early demo on X (opens x.com)
    • Gemma 4 12B
    • Mimi codec
    • RQ depth transformer
    • Synthetic dialogue
    • MLX
    • Modal B200
  2. 02 2026 Real-timeEmbodied RL

    A vision-language drone pilot

    The same problem as full-duplex speech, with a body: a model that has to act on what it sees while the world keeps moving. A Qwen3.5-0.8B pilot watches a 10 Hz camera feed and flies a simulated quadrotor through photoreal Gaussian-splat scans. It learns without any human flight data, from about ten hours of clean synthetic flight plus a GRPO variant that credits each decision against sibling rollouts, so the whole thing trains on a MacBook.

    • Qwen3.5-0.8B
    • MLX
    • GRPO
    • Behaviour cloning
    • Gaussian splatting
  3. 03 2026 Synthetic dataMid-training

    Mid-training on a constitution

    Giving a model a character that holds without a system prompt. I wrote a constitution, turned it into a bilingual corpus of about 1,200 synthetic documents with several teacher models and a triage pass, and mid-trained Qwen 9B and 27B models on it, building on Model Spec Midtraining (Li et al., 2026). A Jacobian lens let me compare the models' internals before and after.

    • Qwen3.6-27B
    • LoRA continued pretraining
    • Multi-teacher synthesis
    • Jacobian lens
    • B200
  4. 04 2026 Post-trainingInterpretability

    Rewarding models with their own activations

    Training a character with rewards read from inside the model instead of from a judge: GRPO from Qwen3-4B to Qwen3.6-27B, rewarded by SAE features, a linear character direction or an activation manifold. Mostly a map of how these rewards get gamed, and of the signals that warn you before they do.

    • GRPO
    • Sparse autoencoders
    • Activation directions
    • TRL
    • vLLM
  5. 05 2026 Interpretability

    Emotion concepts in open models

    How models represent emotion, and whether you can steer it. An independent replication of Anthropic's emotion-concepts study on Gemma 4 and on Qwen3.5 from 0.8B to 27B: 171 emotion directions from a synthetic story corpus, with confound removal, logit-lens readouts and causal steering, and a record of which effects held up.

    • Gemma 4
    • Qwen3.5
    • Representation reading
    • Steering
    • Logit lens

02 — Applied

Applied work

The enterprise half. Models for banks and large companies, where reliability comes first.

AND EMILI 2025–2026

ÆRA

Italian-first enterprise language model

A 4B model for grounded answers, structured JSON and tool calls, designed for on-prem deployment.

  • Trained on synthetic data from seven custom generators (grounded and long-document QA, extraction, multi-turn tool use, editing, reasoning), filtered by an LLM judge.
  • Built to say when the answer isn't in the context instead of guessing.
  • Distilled into a decision engine (yes/no, choice, score) read straight from answer logits, trained on teacher-checked counterfactual pairs. It scores 84% on the public JevBench split, the best of the published sub-10B entries.
SFTSynthetic dataFunction callingDistillation
Credem 2022–2026

Emily

Production banking assistant

A customer-service assistant for a major Italian bank that handles multi-step requests with tool use and a grounded knowledge base, across thousands of conversations a day.

  • Third place at the AIFIn Award.
ProductionNeurosymbolicPost-training
Research prototype 2025

Ontology

Synthetic agent-trajectory engine

A neuro-symbolic pipeline for agentic training data. Task constraints and tool interactions are encoded in a cyclic policy graph; turn-limited stochastic rollouts sample multi-step trajectories, which are realized as natural dialogue with structured tool-call traces. This separates behavioral coverage from linguistic variation, producing supervised training data grounded in an explicit interaction policy.

  • An LLM builds the policy graph under a strict JSON schema; Pydantic validates it, with a rule-based fallback.
  • Each sampled trajectory becomes its own Pydantic schema, so generation can't skip or reorder steps.
NeurosymbolicSynthetic dataTool useStructured outputs

03 — Notebook

More experiments

Smaller experiments and groundwork.

  1. A streaming runtime for Gemma 4

    Drives Gemma 4 12B as one persistent audio and video stream, with micro-turns and barge-in. This was the groundwork for the full-duplex model.

  2. Teaching a model to name injected emotions

    GRPO that trains Qwen3.5-4B to report which emotion vector has been injected into its residual stream, as a small-scale take on learned introspection.

  3. Memory as a modality

    Can a model carry a conversation's state as a latent input instead of text? Three rounds on Qwen3.5-4B: learned memory slots, a gated cross-attention reader, and training-free reuse of the model's own attention and recurrent state.

04 — About

About

I'm a machine learning researcher and engineer based in Italy. I'm less interested in chasing benchmarks than in building things that don't exist yet, and I judge them by what they feel like to use.

That's why most of my time goes into the data: generating it, reading it, and thinking about how the person on the other end will feel.

Training

PyTorchTRLUnslothLoRAMLX

Serving & compute

vLLMMLXModalB200 / H100 / A100

Interpretability

SAEsSteering vectorsLogit & Jacobian lensActivation patching

How I work

  1. Start from the experienceDecide what it should feel like to use, then work backwards to the data, the objective and the architecture.
  2. Look at the samplesMost failures show up in the outputs long before they show up in a metric.
  3. Controls before claimsA result counts only once it beats a baseline, a random control or an ablation.
  4. Keep the failuresNegative results and retractions go in the notes, not the bin.

05 — Contact

Say hello: info@lorenzoiotti.com

I'm happy to go deeper on any of this, or to show you something running.