Handwritten Music Recognition: From Score Image to MusicXML
Optical Music Recognition project, École Polytechnique · Sep 2025 · with Yuguang Yao
Overview
Optical Music Recognition (OMR) is the problem of teaching a computer to read sheet music — recovering not just the symbols on the page but their musical meaning, so the score can be played back or edited. It is substantially harder than OCR: music notation is a two-dimensional language where a symbol's meaning depends on its spatial context (a notehead's pitch is defined by its position relative to staff lines and the active clef), symbols range from tiny augmentation dots to page-high braces, and multiple voices can share the same horizontal position. Handwritten scores add a further layer of variability.
Building on the system of Yang et al. (2024), we first reproduced their detector-plus-linker approach, then extended it into a complete end-to-end pipeline that outputs ready-to-use MusicXML — files that open and play in MuseScore. Our relation-prediction module reaches a Match+AUC of 0.86 on the MUSCIMA++ handwritten benchmark, and the rule-based assembly stage we added handles multi-staff grouping, polyphonic voice alignment, and tuplet detection.

Remark. The architecture splits the problem into perception (what symbols are where — learned by neural networks) and semantics (what they mean together — reconstructed by explicit musical rules). That division of labour is deliberate: detection and linking are pattern-recognition problems with abundant training signal, while music's grammar — how stems, beams and noteheads combine into timed, pitched notes — is known a priori and better encoded as rules than relearned from 140 pages of data.
Dataset & preprocessing
We work with MUSCIMA++: 140 pages of complex handwritten music yielding 3,487 cropped images, annotated with both symbol bounding boxes (117 classes — noteheads, stems, beams, rests, accidentals, clefs, time signatures…) and the ground-truth notation graph of relationships between them. Three preprocessing steps prepare the images for detection: a U-Net performs pixel-wise “destaffing”, erasing the staff lines that overlap nearly every symbol; full pages are sliced into overlapping 640×640 patches so no symbol is cut at a boundary; and the XML notation-graph annotations are converted to YOLO training format.
Detection & relation prediction
Visual detection (YOLOv12). The perception engine is the latest YOLO architecture, trained for 100 epochs with SGD and Mosaic augmentation at 640-pixel resolution. Compared to earlier versions, YOLOv12 notably improves small-object detection — which is exactly what a score demands, where an augmentation dot or accidental occupies a handful of pixels. The model achieves high mAP on the common symbol classes, giving the later stages a solid foundation.
Relation prediction (MLP linker). Detected symbols are discrete boxes; music arises from their connections — which notehead belongs to which stem, which stem to which beam. We cast this as link prediction on a graph whose nodes are the detected symbols. For each symbol pair, an MLP (3 hidden layers) sees geometric features (relative offsets, normalized coordinates, IoU overlap) concatenated with learned 32-dimensional class embeddings for each symbol — letting it learn that noteheads connect to stems but never to clefs — and outputs the probability of a directed edge. On the test set the linker scores 0.86 Match+AUC, the benchmark metric combining edge precision against the ground-truth graph with ranking quality of the predicted probabilities.
Semantic assembly
The notation graph is still not music. The assembly module we added transforms it into a hierarchical score tree through six rule-based phases: grouping staves into systems (via braces and measure separators), slicing time into measures by detecting and merging barlines, assembling notes from primitives (including creating virtual stems when a detection is missed), tracking state-persistent attributes like clefs and key signatures across measures, detecting tuplets with a cascade of four rules plus a duration sanity check, and finally determining pitch. Pitch comes from pure geometry — the notehead's vertical offset from the nearest staff line, measured in half-spacings:
mapped to a diatonic pitch through clef-specific reference lines and adjusted by any detected accidentals.

Remark. Following a single excerpt down the four panels shows what each stage contributes and why the order matters. After destaffing (second panel) the symbols survive intact while the five-line grid disappears — without this, staff lines would corrupt nearly every bounding box. The detection panel shows the coverage problem YOLO solves: hundreds of boxes per excerpt across wildly different symbol sizes. And the final panel makes the graph-structure idea concrete — the magenta and yellow links are the notation graph, tying noteheads to stems to beams, which is the structure the assembler walks to emit notes with correct durations and voices.

Remark. This is the figure that certifies “end-to-end”: the messy handwritten scan from the first panel above has become a clean, editable, playable digital score — polyphonic voices and tuplets included. Getting here is precisely what separates a full OMR system from a symbol detector with good metrics, and it is the part our assembly module contributes over the reproduced baseline.
Limitations & future work
The linker's benchmark score hides a practical weakness: some of its errors are musically illogical mismatches that a human would never make, and the current rule-based reasoning cannot always repair them — a natural next step is small learned models assisting specific assembly sub-steps. The symbol vocabulary also covers only the essential classes so far. Longer term, we want probabilistic, interactive processing in the spirit of Audiveris: surfacing detection confidence to a user interface where recognition results can be corrected by hand — since this project is ongoing, that is where it is headed.