ADR Studio Suite is a professional application for managing ADR (Automated Dialogue Replacement — post-production looping/dubbing work). It covers the entire workflow: video import, generating dialogue lines (cues) via automatic transcription or subtitle import, text management and translation, multi-take recording with teleprompter, comping of the best takes, and finally export to whichever format the recipient needs — a CSV for editing, a Word script for the actor, or a complete DAW session for the mixing engineer.
The application runs as a native desktop app on macOS and Windows (Electron), with a single timeline shared across every phase of the work: the cues created during transcription are the same ones used for recording and export, with nothing to re-copy by hand between programs.
ADR Studio Suite is available in two license tiers, aimed at different roles in the pipeline:
| Area | Light | Pro |
|---|---|---|
| Video, SRT, EDL, CSV import | Yes | Yes |
| Automatic AI transcription | Yes | Yes |
| Cue management (text, actor, timecode) | Yes | Yes |
| Assisted translation | Yes | Yes |
| CSV / SRT / EDL / Word script export | Yes | Yes |
| Multi-take recording (arm/solo, teleprompter, punch in/out) | No | Yes |
| Comping lanes, rating, split take | No | Yes |
| Mixer with original-audio reference | No | Yes |
| AI voice isolation (Demucs) and voice detection (VAD) | No | Yes |
| AI spatial tagging (framing/position in scene) | No | Yes |
| Full DAW session export (Pro Tools, Logic, Nuendo, Reaper, Fairlight, Pyramix, Audition, AAF, FCPXML) | No | Yes |
| Advanced exports (Advanced CSV, Advanced EDL, Excel Report, Multi-Stem ZIP) | No | Yes |
In short: Light covers the entire script preparation and translation work through to text delivery (CSV/Word). Pro adds everything needed inside the recording booth: microphone, takes, comping, and audio delivery ready for the mixing engineer.
On first launch, if no license is present, the app automatically starts in a 30-day free trial mode with every Pro feature unlocked. The remaining-days countdown is visible under Settings → License.
| Note: once the 30 days run out without an activated license, the app locks with a full-screen block until a valid license is imported. |
|---|
Whoever provides ADR Studio Suite (or your organization's purchasing department) will send you a license file with a .lic extension, generated specifically for the machine you're installing on.
| State | Icon | Meaning |
|---|---|---|
| License Active | 🟢 | A valid license has been imported and verified. Tier (Light/Pro) as purchased. |
| Free trial | 🟡 | No license imported: countdown against the 30-day trial, all Pro features active. |
| Trial expired | 🔴 | The 30-day trial period ended with no license. App locks on launch. |
| License expired | 🔴 | A valid license was present but its expiry date has passed. App locks on launch. |
| License invalid | 🔴 | The .lic file fails verification (signature mismatch, wrong Machine ID, corrupted/tampered file). App locks on launch. |
The license signature is cryptographically verified (Ed25519) on every launch: a hand-edited .lic file — for instance, one with the expiry date pushed back — is recognized as invalid and rejected outright, not silently ignored.
The main window is divided into four areas, all visible together: top bar, left sidebar, video monitor, and center timeline.
Organized into sections, top to bottom:
In the Light edition, the Recording, DAW Export and Export Professional sections aren't shown: the interface only displays what that license can actually use.
Shows the current frame of the imported video, with the dialogue text overlaid when relevant (readable by the actor during the take), plus a mirror of the external-monitor window for the control room, if active.
Below the video monitor, in order: the timeline zoom/toolbar row, the color-coded block timeline (one block per cue, color-coded by actor), the comping lanes (Pro only, when a cue has more than one recorded take), and finally the text/editable cue table, with columns for number, selection, play, start, end, actor, spatial position (Pro), and dialogue text.
In the table, each cue's text color indicates its sync status: green if the text fits the take's duration, red if it's likely too long, blue if it's shorter than needed — an automatic estimate based on average reading speed.
This chapter walks through the typical path of a project from start to delivery, step by step. Each step points to the chapter that covers it in depth.
From the sidebar, New Project to start from scratch, or Load Project to resume saved work. The project is saved to a file that includes the cues, settings, and references to the recorded takes.
Sidebar → Import → Video. The app automatically reads the file's duration and frame rate and uses them as the reference for the whole timeline. See Chapter 5.
Three possible routes, which can also be combined:
See Chapter 6 for details on automatic transcription.
Text, actor name, and timecode for every cue remain editable at any time, directly in the table, in both license tiers. This is where the dialogue writer and translator step in to fix what was auto-generated, align actor names, and resolve any overlaps. See Chapter 5.
Arm the cue you want to record (the Arm button in the table or on the timeline), set pre-roll/post-roll if needed, and press Record. The teleprompter shows the text to the actor during the take. Every repetition creates a new take, numbered in sequence. See Chapter 8.
When a cue has multiple takes, the comping lanes let you listen to them, assign a rating, mark the best one (Best Take), and cut/join portions from different takes into a single composite performance, without editing outside the app. See Chapter 9.
Depending on the recipient: CSV or Word Script for whoever works on the text, or a complete DAW session (with the best takes' audio stems already positioned on the timeline) for the mixing engineer. See Chapter 13.
The cue is ADR Studio's basic unit of work: it represents a single line of dialogue, with a start timecode, an end timecode, an assigned actor, and the text to be dubbed. The entire application — transcription, translation, recording, export — revolves around cues.
In the cue table, the Text cell is directly editable: a click places the cursor, and you type as normal. The Actor cell is an autocomplete field: as you start typing, it suggests names already used in the project, and automatically assigns a consistent color to each actor (the same color recurs in the timeline, the comping lanes, and exports).
The Start and End cells are edited by clicking on them: a timecode editor opens in HH:MM:SS:FF format, based on the project's FPS.
The search field above the table (Search by actor name...) filters the cues to show only those for the searched actor — useful in productions with many characters, to focus on one voice actor at a time during a recording session.
Selecting multiple cues via the checkboxes in the table's first column (or Select All Cues from the Cue menu) lets you merge them into a single line with Join Selected Cues (Join) — useful when automatic transcription has split a single sentence into multiple segments.
The ✕ at the end of a row deletes that single cue; the ✕ in the table header (Delete All Cues) clears the entire project — with a confirmation prompt, since this is a destructive action.
| Color | Meaning |
|---|---|
| Green | The text fits within the assigned take's duration |
| Red | The text is likely too long for the available time |
| Blue | The text is shorter than necessary for the available time |
The estimate is based on an average reading speed per language and is meant as a quick heads-up, not an exact measurement: it's normal for some cues to stay "red" by stylistic choice (e.g. rushed lines) — the mixing engineer and dialogue writer always have the final word.
Every change to text, actor, or timecode is recorded in the project history: Ctrl+Z (Cmd+Z on Mac) undoes, Ctrl+Y redoes, for up to 50 steps back.
ADR Studio Suite includes a local transcription engine (based on WhisperX) that listens to the imported video's audio and automatically generates cues: it detects where each line starts and ends, transcribes the text, and estimates the speaker.
Besides standard transcription, a "Studio Transcription" mode is available (menu Transcription → Studio Transcription), aimed at longer sessions or higher quality requirements, with a dedicated local engine and the option to abort partway through if needed (Abort Studio Transcription).
The generated text goes through post-processing that removes typical automatic-transcription "hallucinations" (anomalous repetitions, text generated during silence) before being turned into cues — but a human re-read is still recommended before moving on.
If you already have an SRT file or an ADR CSV prepared elsewhere, you can import it directly (Sidebar → Import) instead of starting from transcription — useful when the script arrives already prepared by another department.
The Text Translation section in the sidebar lets you batch-translate every cue in the project from a source language to a target one.
| Note: automatic translation is a starting point to speed up the dialogue writer/adapter's work — it doesn't replace the lip-sync and rhythm adaptation that dubbing requires. |
|---|
The recording module turns ADR Studio into a genuine ADR booth workstation: per-character arm/solo, configurable pre-roll and post-roll, teleprompter, punch in/out, markers, and level monitoring.
The Arm button in the table (or on the timeline) readies a cue for recording. Shift+Click on Arm arms an entire range of consecutive cues, handy for recording a sequence of lines without stopping to arm each one individually.
The Solo button isolates a cue for listening/playback, muting the others — convenient for focusing on a single line during rehearsal.
Pre-roll is the amount of time (in seconds) played before the cue starts, giving the actor time to settle into the scene's rhythm; post-roll is the time after it ends. These are set per cue or as project-wide defaults in Settings.
Every recording can have a short fade at the start and end (in milliseconds), useful to avoid audio clicks at the beginning/end of a take.
Displays the current cue's text full-screen (or on the external monitor) during the take, with adjustable font size in Settings — designed to stay readable even on a second screen positioned away from the actor.
Lets you re-record only a portion of an existing take, without repeating the entire line from scratch.
Pressing M during playback drops a marker on the timeline, visible in the sidebar's Markers section — useful for flagging a spot to revisit without stopping work.
Tools → External Monitor opens a dedicated window to move onto a second screen in the control room or booth, showing the video and overlaid dialogue text, independent of the main window.
Besides the per-cue mode, a continuous recording mode is available that follows the video without stopping at every line, for sessions with experienced actors who prefer a more natural flow.
Every time an already-armed cue is recorded, a new take is created, numbered in sequence (T1, T2, T3...). When a cue has multiple takes, comping comes into play: choosing, joining, or trimming the best parts of each.
Below the main timeline, the comping lanes show every take for an actor side by side, with direct drag/trim/selection on the waveform: drag to select the portion you want from one take, then move to another lane for the next line.
Every take can receive a rating (stars) and an identifying color, so the best takes stand out at a glance without having to re-listen to all of them every time.
One take per cue can be marked as the Best Take: that's the one the app uses by default for exports (audio stems, DAW session) unless a different choice is made manually during comping.
A take that's too long, or with a mistake partway through, can be split in two (Split) at the exact chosen point, so the final version can combine the good part of one take with another.
Sidebar → Recording → Take Manager opens a window with the full list of takes per cue, to rename, delete, re-listen, and approve — including all cues in the project at once (Approve Takes for All Cues).
Sidebar → Recording → Take Report generates a text summary of every take recorded in the project: how many per cue, which one was approved — useful as a session log to keep or share with production.
The mixer controls three independent volumes during listening and recording: the original video's audio, playback of already-recorded takes, and the microphone input.
Separates the voice from the score/background noise in the video's original audio, using a local AI model (Demucs). Useful for a cleaner, "dialogue-only" waveform during sync work, and for comparing the original line's rhythm against the new recording more easily.
Automatically detects the segments where speech is present in the audio, as a cross-check against the transcription's own segmentation: if speech recognition got a line's alignment wrong, VAD helps pinpoint where the speech actually starts and ends.
Detection-sensitivity presets are available (Settings → AI section), adjustable based on the audio type (clean dialogue, noisy scene, presence of music).
The Spatial Awareness module automatically classifies the framing and position of the speaking character for each cue, by analyzing the video frame at that line's timecode — not just the text, since "on/off screen, back to camera, wide shot" are visual facts that the transcribed text alone can't provide.
| Code | Meaning |
|---|---|
| FC | Off Screen |
| IC | On Screen |
| DS | Back to Camera |
| SD | Lying Down |
| ACC | Crouched |
| MOV | In Motion |
| PP | Close-Up |
| CL | Wide Shot |
| CLL | Extreme Wide Shot |
| SM | Silent/Mute |
The assigned tag appears in the Pos. column of the cue table, and can be edited by hand at any time if the automatic classification isn't correct.
Classification runs on a local vision+language model (Qwen2.5-VL, GGUF format, run through llama-cpp-python), completely offline: no frame ever leaves the computer. Since it's a 10-category classification task with constrained output, not free text generation, it performs well on CPU alone, with no dedicated graphics card required.
Menu Transcription → Retry Missing Spatial Tags reprocesses only the cues that don't yet have a tag assigned (for example, after manually adding new cues), without having to rerun classification on the whole project.
This feature requires both model files (language + vision projector) to be present in the application's Models folder. If they're missing or incompatible, classification will error out — see Chapter 16, "Spatial Awareness isn't working."
ADR Studio Suite exports directly into whatever format the recipient uses, avoiding manual intermediate steps and possible sync loss.
| Format | What it's for |
|---|---|
| ADR CSV | Spreadsheet-format cue list, for importing into other software or tabular review |
| SRT | Standard subtitle file, compatible with most video editors |
| Word Script | A print-ready .docx document, to hand to the actor or dubbing director |
| EDL | Edit Decision List, for editing |
Generates a session ready to open directly in the chosen DAW, with one track per character and the audio files for each take (normally the Best Take) already positioned at the correct timecodes:
| DAW / format | Notes |
|---|---|
| Pro Tools (AAF) | An .aaf session importable via File > Import > Session Data |
| Logic | Native Logic session |
| Nuendo | Session with AAF support |
| Reaper | .rpp project, one track per character |
| Fairlight | Session with AAF and FCPXML support |
| Pyramix | Dedicated session |
| Adobe Audition | Dedicated session |
| FCPXML | For editing, with absolute media references |
The Settings panel (Cmd/Ctrl+, or Tools → Settings...) gathers every project- and application-level option.
Interface language (Italian/English) and default source/target languages for assisted translation.
The Reset AI Settings button restores this section to its defaults — useful if a custom configuration has stopped working as expected.
Current status, Machine ID, and the import button — see Chapter 2 for details.
| Key(s) | Action |
|---|---|
| Click on the timeline | Move the playhead (seek) to that point |
| Shift + Click | Create a new cue at the clicked point |
| M | Drop a marker at the current timecode |
| Shift + L | Toggle loop |
| R | Start/stop recording |
| Spacebar | Play / Pause |
| J / K / L | Shuttle: rewind / pause / forward (variable speed on hold) |
| Ctrl/Cmd + Z | Undo |
| Ctrl/Cmd + Y | Redo |
| Ctrl/Cmd + N | New Project |
| Ctrl/Cmd + O | Open Project |
| Ctrl/Cmd + S | Save Project |
| Ctrl/Cmd + Shift + E | Export to DAW... |
| Ctrl/Cmd + , | Open Settings |
| Shift + Click on Arm | Arm a range of consecutive cues |
This chapter collects the most common issues encountered in day-to-day use and during installation, along with their solutions.
The app locks intentionally in these three states — by design, not by error: it's confirmation that a valid .lic file is needed. Check that you've imported the latest file you received (Settings → License) and that the Machine ID shown matches the one given to whoever generated the license. If the file was modified even slightly (for instance with a text editor), its signature is no longer valid and a new file must be requested.
This means the internal Python environment used for local AI is either missing the required module or can't find the model files. It's a known issue especially on manually compiled or custom builds:
The notes below concern whoever builds the executable from source code, not day-to-day use of the already-installed app.
| Feature | Light | Pro |
|---|---|---|
| Video / SRT / EDL / CSV import | Yes | Yes |
| Automatic AI transcription | Yes | Yes |
| Cue editing (text, actor, timecode) | Yes | Yes |
| Assisted translation | Yes | Yes |
| CSV / SRT / EDL / Word Script export | Yes | Yes |
| Recording (arm/solo, teleprompter, punch in/out) | — | Yes |
| Comping, rating, split take | — | Yes |
| Mixer with original-audio reference | — | Yes |
| Voice isolation (Demucs) and VAD | — | Yes |
| Spatial Awareness AI | — | Yes |
| DAW session export (Pro Tools, Logic, Nuendo, Reaper, Fairlight, Pyramix, Audition) | — | Yes |
| Export Professional (Multi-Stem ZIP, Advanced CSV/EDL, Excel Report) | — | Yes |
| Term | Meaning |
|---|---|
| ADR | Automated Dialogue Replacement — recording dialogue in sync with picture, in a booth, to replace or supplement the original production audio |
| Cue | A single line of dialogue, with start/end timecode, actor, and text |
| Take | A single recording/repetition of a cue |
| Comping | Assembling a cue's final version by choosing/joining the best parts from multiple takes |
| Best Take | The take marked as best for a cue, used by default in exports |
| Arm | Readying a cue for recording |
| Pre-roll / Post-roll | Listening time before and after the cue respectively, to give rhythmic context |
| Punch in/out | Re-recording only a portion of an existing take |
| Teleprompter | On-screen display of the line's text during the take |
| VAD | Voice Activity Detection — automatic detection of speech segments in an audio signal |
| Stem | An isolated audio file for a single source (e.g. one character), ready for mixing |
| EDL | Edit Decision List — a standard list of editing changes, portable across software |
| AAF | Advanced Authoring Format — a session exchange format between professional software (e.g. Pro Tools) |
| Timecode | A time reference in HH:MM:SS:FF format (hours:minutes:seconds:frames) |
| FPS | Frames per second of the video, the basis for every timecode calculation in the project |