Skip to content

Pick a time for your free call

Loading calendar…

Calendar not loading? Open it on Cal.com

Timeline

  1. 01 The Setup Handing Over the Mouse
    1. Why I tried it
    2. One model, four tools
  2. 02 The Process Every Song
    1. One song, start to finish
    2. Where the mouse mattered
  3. 03 The Results What Survived
    1. The ears problem
    2. What survived
  4. 04 The Lessons Back to Me
    1. What it taught me
    2. Two days, eleven mixes
    3. If you work in a creative tool
← Back to Blog

Build log 13 min read 4 phases

I Let an AI Drive Pro Tools for an Entire Record (The Best Thing It Did Was Put Things Back)

Eleven mixes, two days, one AI model driving real Pro Tools with a mouse and a scripting bridge. It fixed a mix that was clipping at +9 dBTP, trimmed every song to a shared headroom target, tested at least fourteen tonal ideas, and put every one of them back. Here's exactly how it worked, and what it taught me about my own mixing.

TLDR

I’ve been engineering audio for over fifteen years, and I just let an AI drive Pro Tools for the final mix pass on an eleven-song record. It ran on GPT-6 Astra inside Codex, on the Ultra reasoning setting, using computer use to click through Pro Tools like a person would, plus a scripting bridge, local loudness meters, and a second AI model for listening. Over two days it fixed a mix that was clipping at +9 dBTP, brought every song into a shared headroom range, kept one drum revision I approved by ear, and put every other tonal idea it tested, at least fourteen of them, back exactly how it found them. The most impressive thing it did was know what not to keep.


Why I Handed Over the Mouse

The mixes existed. They were mine. The problem was the thing that always happens with a record made over time: every song was mixed in a different week, in a different mood, at a different level. Listen to them back to back and the record doesn’t sit together.

So the brief wasn’t “make it sound better.” It was consistency without flattening. Here’s how it went into the project notes: “similar mastering headroom, familiar frequency balance and vocal presentation where useful, with precise changes that preserve every track’s soul, vibe, aesthetic and character.”

The notes are explicit that it was a premaster pass, “not a request to force identical spectra, LUFS, processing or density.” I didn’t want eleven songs that sound the same. I wanted eleven songs that sound like they belong on the same record.

That’s a lot of careful, repetitive work across dozens of tracks and hundreds of plugin settings per song. Which is exactly the kind of work I wanted to see if an AI could do without wrecking anything.


The Setup: One Model, Four Tools

This is the part most people get wrong when they hear “AI mixed it,” so I want to be precise. It wasn’t one tool. It was one model running four tools, and only one of those tools could hear.

  1. A scripting bridge. Avid has an official scripting layer for Pro Tools. We built a small bridge to it that ran locally on the Mac. It could read the full track list and plugin inventory and render the final stereo file.
  2. Computer use. Codex has a built-in computer use feature that sees the screen and operates the Mac through the accessibility system: moving the mouse, clicking, typing. This handled everything the scripting bridge couldn’t, which turned out to be more than I expected.
  3. Local meters. A separate measurement tool ran an industry-standard loudness meter, with a second independent library as a cross-check, plus spectrum and stereo correlation analysis on every render. When the system says a song sits at -21.9 LUFS, that’s a measurement, the same one a meter plugin gives me.
  4. A second AI for listening. Short, loudness-matched excerpts went to Google’s Gemini for a plain-language opinion on how something sounds. I pre-approved that for any track I hand over for a mix.

The model running the session and making every change was GPT-6 Astra. The model that ever “listened” to anything was a completely different one. That distinction ends up mattering a lot.

Before it touched a single setting, it scanned my actual plugin library instead of assuming a generic rig: 358 plugin bundles across four formats, including the exact versions of the EQs, compressors, and reverbs I actually use. Nearly every decision on the record used tools already in that scan. It didn’t bring in anything new.


One Song, Start to Finish

Every song ran through the same routine. I opened a fresh task, attached the song’s live Pro Tools session, and gave it one line. I’d set a standing rule early: “Whenever I invoke this skill it means that I would like you to mix the track whose session file I am giving you.” No back-and-forth about scope.

Then:

  1. Protect the original first. Before any change, it fingerprinted the session file on disk, kept a backup, and did a Save As to a new, clearly named copy. Every delivered song’s notes confirm the original file is byte-for-byte unchanged.
  2. Find the real ending. I leave practice takes on the timeline after the song ends. Every pass had to find the actual musical ending and make sure none of the extra material ended up in the bounce.
  3. Read everything before touching anything. Every active track’s routing, sends, pans, automation state, and every plugin’s exact parameter values. Not “there’s a compressor on the vocal.” The literal threshold, ratio, attack, release, and mix.
  4. Measure the untouched baseline. A fresh bounce of the mix as I left it, measured for loudness, true peak, and loudness range, before any change.
  5. Test one change at a time. For each idea, a level trim, a filter, a compressor setting, it made exactly one change, rendered a short loudness-matched excerpt, and prepared an identical, unchanged copy as a hidden control. The pair went out for a blind listening opinion while the plugin’s own meters gave an objective number.
  6. No proven benefit, no kept change. If the test came back unclear, contradictory, or unsupported, the setting went back to its exact original value and got logged as “tested, not proven.”
  7. Render, re-verify, audit itself. Final stereo bounce, measured again, then a separate check that every track, plugin, and automation lane it didn’t mean to touch was still identical.
  8. File it where I told it to. Every bounce went into that song’s own bounced-files folder, under a dated subfolder, never overwriting an earlier version.
  9. Teach me. Each song got a readable notes file for me, plus a machine-readable record that tagged every finding as a verified problem, an untested hypothesis, or an intentional choice to preserve.

That last step was my idea, and it’s the one I’d tell anyone to steal. I wanted to learn what I do in my own sessions that might not be ideal. But I didn’t want a single song’s quirk turning into a “pattern.” So every failed test stayed in the record, clearly labeled, instead of quietly disappearing.


Where the Mouse Mattered

I expected the scripting bridge to do most of the work and computer use to handle the odd popup. It was closer to the other way around.

The scripted, official way to do a Save As crashed Pro Tools with an internal assertion error. Every time. So the most basic protective step in the whole routine, saving a new copy before touching anything, ended up being done by computer use: the AI opening the real File menu and clicking through the dialog, exactly like I would.

Computer use also dismissed missing-plugin warnings, double-clicked into plugin fields, and typed exact values, and it read the plugin windows directly, because the scripting layer on my version of Pro Tools couldn’t read automation breakpoints at all.

It wasn’t perfect, and the misses taught me the most:

  • Once, a click meant to open a plugin’s parameter editor fired a different action and quietly set a reverb’s dry/wet mix to 100%. It caught that immediately and restored it.
  • On one specific plugin, double-clicking a control did the opposite of what the same gesture does on every other plugin: instead of opening the field for typing, it reset the value to zero.

The rule that came out of it is simple: never trust that a click worked. Check the result of every click, every time. Plugin vendors don’t agree on what a double-click means.


The Ears Problem

Here’s the part I didn’t see coming. The weakest link wasn’t the mouse. It was the ears.

The listening model was wrong often, and in specific, checkable ways:

  • It described two test renders as identical when they were measurably different.
  • It called one passage quieter when the meters said it was louder.
  • It made claims about the arrangement that weren’t true, like saying there was no kick or bass in a section that had both.
  • In one case it preferred a change in a longer comparison, then couldn’t tell the two versions apart in a shorter blind version of the exact same test, on the same song, the same day.
  • The service itself failed constantly: quota errors and outages on almost every song, including a full-record listening check near the end that never came back.

The system’s own written policy, which I now agree with completely: “Confident prose is not measurement or proof of master readiness.”

And the line from its own notes on the final song that I respect the most: “I have no direct personal hearing, and there is no successful exact-final full-song AI verdict.”

That’s an AI telling me what it can’t do, in writing, at the end of a long job. That’s worth more to me than a confident paragraph about warmth and punch.


What Survived

This is the headline, and it’s not the one people expect.

Across eleven delivered mixes, the only changes that survived into the final files were output-level and headroom moves (trimming the master and, on most songs, bypassing the limiter) plus exactly one creative revision: a drum EQ and compression pass on one song. Every other tonal or dynamics idea it tested, at least fourteen of them, got tried, measured, and put back exactly as it was.

That one drum revision made it in because I listened to it and said, “I really like this mix!” Not because a model said so. In fact, across the whole record, the AI’s opinion was never by itself the reason a tonal change was kept. Every change that stayed was either an objective fix or something I approved by ear.

And there were real fixes. The first song’s original bounce was hitting +9.01 dBTP, well over full scale. On that same song it found a routing problem: two separate output paths, and only one of them had been turned down. That’s the kind of thing you stop hearing after the hundredth playback, and the kind of thing you find by reading a session instead of listening to it.

Where everything landed:

MixLoudnessTrue peakLoudness range
1-25.9 LUFS-5.80 dBTP6.8 LU
1, acoustic version-21.8 LUFS-3.10 dBTP8.6 LU
2-21.3 LUFS-3.39 dBTP8.6 LU
3-22.9 LUFS-4.60 dBTP4.7 LU
4-20.9 LUFS-3.67 dBTP18.8 LU
5-22.9 LUFS-3.80 dBTP5.4 LU
6-23.4 LUFS-4.99 dBTP10.2 LU
7-24.3 LUFS-5.60 dBTP8.7 LU
8-21.3 LUFS-4.70 dBTP9.3 LU
9-20.0 LUFS-4.56 dBTP6.5 LU
10-21.9 LUFS-4.80 dBTP4.2 LU

All eleven are 48 kHz, 32-bit float, with zero clipped or invalid samples. True-peak headroom lands between 3.1 and 5.8 dB across the record, and that range came from the two mixes I’d already approved, not from a number picked in advance. Loudness still varies by 5.9 LU from the quietest to the loudest song, on purpose. This was never a leveling pass. Mastering is its own step.


What It Taught Me About My Own Mixes

The last song came with a bigger ask: mix it, then go back through all ten previous songs’ evidence and tell me what I actually do wrong, with a warning not to invent a pattern out of one example.

The opening line of that review is my favorite sentence from the whole project: “Your most consistently useful next habit is to decide the mix’s delivery level deliberately and verify the complete output. The evidence does not establish a recurring audible compressor, EQ or reverb mistake.”

In other words: no recurring EQ, compression, or reverb mistake showed up across ten songs. The habit to fix was my output stage. That’s a humbling and weirdly comforting thing to learn after fifteen years.

The five habits it pulled out, which I’ve put on a sticky note:

  1. Decide the delivery stage on purpose, and check every output path, before touching the overall level.
  2. Write down the specific, timestamped problem before turning any knob. “Mud” is a question to test, not an automatic cut.
  3. Trace the whole signal path first. Shared recordings, which bus a track actually feeds, which way a dynamic EQ is moving. Several settings that looked suspicious in isolation made sense in context.
  4. Always compare against an unchanged copy, and keep the failed tests in the notes.
  5. Judge tone at matched volume, then judge the running order at real volume, separately. They answer different questions. That practice comes from Ian Stewart’s published approach to album mastering.

Two Days, Eleven Mixes

The first seven mixes happened on September 8, starting at 9:46 in the morning. The last four, plus the full cross-song review, happened on September 9. One song’s task window stayed open about eight hours, including long waits on the listening service. The final song, including the review of the entire record, took about an hour and fifteen minutes.

To be clear about where this stands: this was the premaster mix pass. When it finished, I’d approved two of the eleven mixes, and the rest were candidates waiting on my ears. Mastering is its own step after that. The AI did the careful, measured, repetitive part. The judgment is still mine, and I think that’s exactly the right split.


What This Means If You Work in a Creative Tool

This isn’t really a music story. It’s a story about computer use.

For most of software history, if you wanted to automate a professional tool, you needed an API, and most pro tools don’t have a good one. Now any app with a screen can be driven. Pro Tools. A video editor. The design tool your team uses every day. The rules that made this work would apply to any of them:

  1. Protect the original before anything else. Automatically, every time, and verify it afterward.
  2. Measure with real instruments. A model’s description of a result is not the result.
  3. Change one thing at a time, against a control. And keep the failures in writing.
  4. Verify every click. The interface will surprise you.
  5. Keep the taste call human. Let the AI do the careful work, and keep the last word for the person whose name is on it.

The best thing the AI did on this record was put things back. I’ve worked with human engineers who couldn’t do that.


I appreciate you reading through this one. It sits right where my two worlds meet. If you want to talk about putting AI to work inside a tool your business already runs on, book a call from the home page, or check out the YouTube channel @generussai.

Build first, learn fast.

Keep reading