Home
← Projects

Face-Mask Generation with Stable Diffusion Inpainting

Generative modelling course, École Polytechnique · 2026 · with Mohamed Amine Amrani

Report (PDF)Poster (PDF)GitHub repository

Overview

Face-recognition systems trained on unmasked faces often fail on masked ones — the occlusion removes half the facial features — yet public masked-face datasets remain scarce, lack diversity, and rarely annotate correct vs incorrect mask wearing. Collecting more real face data is costly and privacy-sensitive, so we asked the generative question instead: given a clean portrait, can we synthesise a photo-realistic masked version of the same person?

We built a two-stage pipeline — a U-Net that predicts where the mask should go, then a diffusion model that generates it — and compared two diffusion approaches against the strongest GAN baseline (CycleGAN): a denoising diffusion probabilistic model (DDPM) trained from scratch, and a pretrained Stable Diffusion inpainting model fine-tuned with LoRA. The LoRA-adapted Stable Diffusion won clearly, reaching FID 17.69 against 28.48 for CycleGAN. Training data combined 70k clean faces from FFHQ with the 133k synthetic masked faces of MaskedFace-Net.

Pipeline diagram: a clean face goes through a U-Net that outputs a predicted mask region; the mask is combined with the clean face to define the editable area; a diffusion model iteratively denoises this input to produce the generated masked face.
The two-stage pipeline: the U-Net localises the mask region, the diffusion model synthesises the mask inside it.

Remark. The decoupling in this diagram is the design decision that matters. Because generation is conditioned on an explicit binary mask, the diffusion model only ever edits the nose-and-mouth region — the rest of the portrait passes through untouched, which is what preserves identity. It also makes the system controllable at inference: hand it a different mask region and it will follow, which lets it generalise to off-distribution photos where an end-to-end image-to-image translator like CycleGAN degrades.

Stage 1 · U-Net mask predictor

The first stage learns to segment the facial region a mask should cover, from 2,000 training pairs of clean faces and binary mask annotations at 128×128 resolution (1,000 validation, 1,000 test). The U-Net is trained for 30 epochs with AdamW, mixed precision on a single T4 GPU, and a composite loss balancing binary cross-entropy with a soft-Dice term that fights the class imbalance (the mask occupies a small fraction of the pixels):

Lmask=0.5 LBCE+0.5 LDice\mathcal{L}_{\text{mask}} = 0.5\,\mathcal{L}_{\text{BCE}} + 0.5\,\mathcal{L}_{\text{Dice}}
Six test faces of diverse ages and ethnicities on the top row, each with its predicted binary mask below: green mask-shaped regions covering nose, mouth and chin, with IoU scores between 0.92 and 0.95.
Predicted mask regions on test faces (green), with per-image IoU scores of 0.92–0.95.
Two histograms over the test set: IoU scores concentrated near 0.93 with mean 0.927 and median 0.942, and Dice scores concentrated near 0.96 with mean 0.962 and median 0.970.
Distribution of IoU (mean 0.927) and Dice (mean 0.962) over the 1,000-image test set.

Remark. The predictor is reliable and consistent, not just good on average. The qualitative panel shows it locking onto the nose-to-chin region across ages, skin tones, glasses, and slight head turns. The histograms add the statistical version of that statement: both distributions are tight and left-skewed, with medians (0.942 IoU, 0.970 Dice) above the means — meaning most predictions are excellent and errors come from a thin tail of hard cases rather than broad mediocrity. That matters downstream: the diffusion stage edits exactly the region this network predicts, so localisation errors would become visible artifacts in the generated faces.

Stage 2 · Two diffusion generators

DDPM from scratch. A conditional diffusion model (U-Net backbone with sinusoidal time embeddings) is trained to reconstruct the masked face given the clean image with the mask region blanked out, plus the binary mask itself, concatenated as conditioning channels. It uses a cosine noise schedule over 250 timesteps and DDIM sampling with 100 steps at inference to cut runtime. Trained on the same 2,000 images as the mask predictor.

Stable Diffusion inpainting + LoRA. The second branch adapts the pretrained runwayml/stable-diffusion-inpainting model. The tokenizer, CLIP text encoder, VAE, and original U-Net weights all stay frozen; adaptation happens through rank-16 LoRA modules inserted broadly into the U-Net (attention projections plus additional projection and convolution layers). Training ran 30 epochs at 256×256 on aligned FFHQ–MaskedFace-Net pairs with the fixed prompt “a high quality portrait photo of a person wearing a protective face mask”; inference uses 30 denoising steps with classifier-free guidance 7.5.

Diagram of the Stable Diffusion branch: a clean face and its mask region feed into the Stable Diffusion inpainting model; LoRA adapters, shown as a small side module, inject trainable low-rank updates into the frozen inpainting U-Net, which outputs the generated masked face.
The Stable Diffusion branch: the pretrained inpainting model stays frozen; only the small LoRA adapters (green) are trained.

Remark. The visual proportions of this diagram are honest about where the knowledge lives: the big frozen inpainting stack carries the pretrained image prior, while the LoRA adapters — the small yellow module on the side — are the only part that learns the mask-generation task. That asymmetry is the whole economics of the approach: adapting a large model on a single GPU by training a tiny fraction of its parameters, while the mask region still tells the model exactly which pixels it is allowed to change.

Seven panels showing the denoising trajectory: the clean input face, then pure noise at step 1, gradually resolving through steps 5, 10, 15 and 19 into the final generated image of the same man wearing a blue surgical mask.
The reverse diffusion trajectory of the Stable Diffusion branch over 20 steps: from noise to a masked portrait.

Remark. Reading the strip left to right shows diffusion's coarse-to-fine character: at step 10 the global structure is already decided — head pose, skin tone, the blue mask's placement — while the last steps only sharpen texture like the mask's pleats. Notice also how strongly the final image resembles the input face in hair, ears, and background: the pretrained latent prior reconstructs the unedited regions almost exactly, which is precisely the identity-preservation behaviour the metrics later confirm (background MAE three times lower than mask-region MAE).

Results

MethodEpochsTrain FID ↓Test FID ↓
CycleGAN (baseline)3028.4328.48
DDPM (from scratch)5047.0948.93
Stable Diffusion + LoRA30—17.69

The fine-tuned Stable Diffusion branch also scores well on the paired-image metrics against ground-truth masked faces:

FID (lower is better)17.69
LPIPS (lower is better)0.073
PSNR (higher is better)23.56 dB
MAE — whole image0.0362
MAE — mask region0.0802
MAE — background0.0268
A grid comparing four test faces across three methods: CycleGAN masks look pasted-on and sometimes misaligned, DDPM masks are blurry and wash out surrounding detail, while the Stable Diffusion fine-tuning row shows sharp, naturally fitted masks with faces and backgrounds intact.
The same four test faces masked by CycleGAN (top), DDPM (middle), and LoRA-fine-tuned Stable Diffusion (bottom).

Remark. The grid makes the numbers tangible. CycleGAN's masks sit on the face like stickers — plausible when the input resembles its training distribution, but the mask geometry drifts on the profile and sunglasses cases. The DDPM row shows what training a diffusion model from scratch on only 2,000 images buys: masks appear in the right place (the predictor works) but are blurry, and skin texture around them degrades — hence its FID of ~49. The Stable Diffusion row is the payoff of transfer learning: crisp pleats, correct wrapping around the chin, glasses and hair untouched. Interestingly, at the poster stage an 8-epoch LoRA run still trailed CycleGAN (FID ≈ 33 vs 28); the final 30-epoch run more than closed that gap — the ranking between methods was itself a function of training budget.

Limitations & takeaways

The headline result is that parameter-efficient adaptation beats both alternatives: with the backbone frozen and only rank-16 LoRA updates trained, the pretrained latent prior contributes realism that neither a from-scratch DDPM (too little data) nor a GAN (brittle off-distribution) could match. But the metrics don't tell the whole story: some generated faces remain visibly distorted even when FID, LPIPS, and PSNR look good — global statistics can miss local facial damage, and none of these metrics truly measures identity preservation. The conclusion we drew is to lean harder on the pipeline's own structure: constraining generation to the explicitly predicted mask region (as the DDPM branch does) is the safer path, and future work should combine that hard constraint with the pretrained prior of Stable Diffusion — plus evaluation protocols that measure identity directly.