AI MIDI System
AI · Audio · 2024 — 2025
The audio half an ML team didn't have, turning quantized scores into human-sounding piano
R&R Records had the data and the machine learning, and no one who knew how music is actually processed. Edge Audio Labs joined the team and built the music-and-audio engine that turns a quantized score into a performance a listener would take for a real player.
- C++
- PyTorch
- AWS
- Client
- R&R Records
- Industry
- AI · Audio
- Engagement
- 2024 — 2025
- Deployment
- AI model / engine
- Problem: R&R Records had a strong data and ML team and no audio or MIDI expertise to make piano scores sound humanly performed.
- What we built: the music-and-audio engine of a production pipeline that interprets piano scores and renders them, combining three AI models.
- Result: quantized scores become distinct, human-sounding performances, in a pipeline built to render at scale.
The Pain
AI Pianist came to Edge Audio Labs as staff augmentation: R&R Records, a company producing background music for video, strong on machine learning and with nobody who knew audio or MIDI. Getting a machine to perform a score, rather than replay it at a fixed velocity, is a music problem, and it was the gap in their stack.
- Quantized isn't played. Notes sitting on a grid read as a machine, whatever instrument renders them.
- The models alone fell short. Performance models off the shelf didn't produce takes that held up.
- Nobody in-house could close it. Strong on data and ML, no one who worked in audio or MIDI.
What We Did
Two decisions shaped the engine.
- Humanization as a music problem, not a modeling one. More than one performance model shapes each take, and on top of them Edge Audio Labs built the timing, dynamics and phrasing logic that carries most of the realism.
- Score every render automatically. So the pipeline could run unattended instead of needing someone to listen to each take.
What We Built
- Score interpretation. Reads MusicXML and MIDI scores and prepares them to be performed, not just played back.
- Humanization. Turns rigid, quantized notes into a take that sounds played by a human, not sequenced.
- Three models combined. Combines three AI models to shape each interpretation, rather than leaning on a single one.
- Rendering engine. Renders the interpreted performance to audio and applies effects, in C++.
- Evaluation. Scores every render so only convincing performances move forward.
- Runs on AWS. The whole flow runs unattended and at scale in the cloud.
Quantized in, Performed Out
The same bars, before and after the engine: the grid on the left, the performance on the right, with timing, dynamics and phrasing shaped by the models.


Where It Got Hard
In the definition.
Nobody can write down what makes a performance sound human. You know a take is wrong the moment you hear it, and that's no use to a pipeline meant to run without anyone listening.
Specs
- DeploymentAI model / engine
- DomainSolo piano
- InputMusicXMLMIDI scores
- OutputRendered audio performances
- ModelsThree AI models, in combination
- StackC++PyTorchAWS
- Use caseBackground music for video
The Value
R&R Records got the audio half they didn't have, without hiring for it.
The missing half, without a hiring round
MIDI, performance logic and native rendering, delivered as part of their team.
Nobody has to listen to every take
Only convincing performances move forward — the pipeline runs unattended.
Music AI, cloud and C++ in one place
Music AI, cloud data engineering and native rendering, working as one system.
Scores in, performances out.
How We Worked
Staff augmentation. Edge Audio Labs joined the client's team and owned the audio half of the pipeline: music AI, MIDI and native rendering.
See how we work




