Face-Mask Generation with Stable Diffusion Inpainting
Generative modelling course, École Polytechnique · 2026 · with Mohamed Amine Amrani
Overview
Face-recognition systems trained on unmasked faces often fail on masked ones — the occlusion removes half the facial features — yet public masked-face datasets remain scarce, lack diversity, and rarely annotate correct vs incorrect mask wearing. Collecting more real face data is costly and privacy-sensitive, so we asked the generative question instead: given a clean portrait, can we synthesise a photo-realistic masked version of the same person?
We built a two-stage pipeline — a U-Net that predicts where the mask should go, then a diffusion model that generates it — and compared two diffusion approaches against the strongest GAN baseline (CycleGAN): a denoising diffusion probabilistic model (DDPM) trained from scratch, and a pretrained Stable Diffusion inpainting model fine-tuned with LoRA. The LoRA-adapted Stable Diffusion won clearly, reaching FID 17.69 against 28.48 for CycleGAN. Training data combined 70k clean faces from FFHQ with the 133k synthetic masked faces of MaskedFace-Net.

Remark. The decoupling in this diagram is the design decision that matters. Because generation is conditioned on an explicit binary mask, the diffusion model only ever edits the nose-and-mouth region — the rest of the portrait passes through untouched, which is what preserves identity. It also makes the system controllable at inference: hand it a different mask region and it will follow, which lets it generalise to off-distribution photos where an end-to-end image-to-image translator like CycleGAN degrades.
Stage 1 · U-Net mask predictor
The first stage learns to segment the facial region a mask should cover, from 2,000 training pairs of clean faces and binary mask annotations at 128×128 resolution (1,000 validation, 1,000 test). The U-Net is trained for 30 epochs with AdamW, mixed precision on a single T4 GPU, and a composite loss balancing binary cross-entropy with a soft-Dice term that fights the class imbalance (the mask occupies a small fraction of the pixels):


Remark. The predictor is reliable and consistent, not just good on average. The qualitative panel shows it locking onto the nose-to-chin region across ages, skin tones, glasses, and slight head turns. The histograms add the statistical version of that statement: both distributions are tight and left-skewed, with medians (0.942 IoU, 0.970 Dice) above the means — meaning most predictions are excellent and errors come from a thin tail of hard cases rather than broad mediocrity. That matters downstream: the diffusion stage edits exactly the region this network predicts, so localisation errors would become visible artifacts in the generated faces.
Stage 2 · Two diffusion generators
DDPM from scratch. A conditional diffusion model (U-Net backbone with sinusoidal time embeddings) is trained to reconstruct the masked face given the clean image with the mask region blanked out, plus the binary mask itself, concatenated as conditioning channels. It uses a cosine noise schedule over 250 timesteps and DDIM sampling with 100 steps at inference to cut runtime. Trained on the same 2,000 images as the mask predictor.
Stable Diffusion inpainting + LoRA. The second branch adapts the pretrained runwayml/stable-diffusion-inpainting model. The tokenizer, CLIP text encoder, VAE, and original U-Net weights all stay frozen; adaptation happens through rank-16 LoRA modules inserted broadly into the U-Net (attention projections plus additional projection and convolution layers). Training ran 30 epochs at 256×256 on aligned FFHQ–MaskedFace-Net pairs with the fixed prompt “a high quality portrait photo of a person wearing a protective face mask”; inference uses 30 denoising steps with classifier-free guidance 7.5.

Remark. The visual proportions of this diagram are honest about where the knowledge lives: the big frozen inpainting stack carries the pretrained image prior, while the LoRA adapters — the small yellow module on the side — are the only part that learns the mask-generation task. That asymmetry is the whole economics of the approach: adapting a large model on a single GPU by training a tiny fraction of its parameters, while the mask region still tells the model exactly which pixels it is allowed to change.

Remark. Reading the strip left to right shows diffusion's coarse-to-fine character: at step 10 the global structure is already decided — head pose, skin tone, the blue mask's placement — while the last steps only sharpen texture like the mask's pleats. Notice also how strongly the final image resembles the input face in hair, ears, and background: the pretrained latent prior reconstructs the unedited regions almost exactly, which is precisely the identity-preservation behaviour the metrics later confirm (background MAE three times lower than mask-region MAE).
Results
| Method | Epochs | Train FID ↓ | Test FID ↓ |
|---|---|---|---|
| CycleGAN (baseline) | 30 | 28.43 | 28.48 |
| DDPM (from scratch) | 50 | 47.09 | 48.93 |
| Stable Diffusion + LoRA | 30 | — | 17.69 |
The fine-tuned Stable Diffusion branch also scores well on the paired-image metrics against ground-truth masked faces:
| FID (lower is better) | 17.69 |
| LPIPS (lower is better) | 0.073 |
| PSNR (higher is better) | 23.56 dB |
| MAE — whole image | 0.0362 |
| MAE — mask region | 0.0802 |
| MAE — background | 0.0268 |

Remark. The grid makes the numbers tangible. CycleGAN's masks sit on the face like stickers — plausible when the input resembles its training distribution, but the mask geometry drifts on the profile and sunglasses cases. The DDPM row shows what training a diffusion model from scratch on only 2,000 images buys: masks appear in the right place (the predictor works) but are blurry, and skin texture around them degrades — hence its FID of ~49. The Stable Diffusion row is the payoff of transfer learning: crisp pleats, correct wrapping around the chin, glasses and hair untouched. Interestingly, at the poster stage an 8-epoch LoRA run still trailed CycleGAN (FID ≈ 33 vs 28); the final 30-epoch run more than closed that gap — the ranking between methods was itself a function of training budget.
Limitations & takeaways
The headline result is that parameter-efficient adaptation beats both alternatives: with the backbone frozen and only rank-16 LoRA updates trained, the pretrained latent prior contributes realism that neither a from-scratch DDPM (too little data) nor a GAN (brittle off-distribution) could match. But the metrics don't tell the whole story: some generated faces remain visibly distorted even when FID, LPIPS, and PSNR look good — global statistics can miss local facial damage, and none of these metrics truly measures identity preservation. The conclusion we drew is to lean harder on the pipeline's own structure: constraining generation to the explicitly predicted mask region (as the DDPM branch does) is the safer path, and future work should combine that hard constraint with the pretrained prior of Stable Diffusion — plus evaluation protocols that measure identity directly.