Automated plankton imaging produces severely long-tailed datasets, where the rare taxa of greatest ecological interest have too few images to train or evaluate classifiers reliably. We generate synthetic plankton imagery conditioned on taxonomy: a CLIP encoder is adapted on a large plankton corpus with a ranked contrastive objective extended to deep, ragged taxonomies, then frozen to condition a parameter-efficient diffusion transformer. We evaluate synthetic sample quality on distributional fidelity and downstream classifier utility.
@inproceedings{ivanova2026plankton,title={Multimodal Taxonomic Conditioning for Generative Plankton Imagery},author={Ivanova, Daniela and Göksu, Özgü and Pugeault, Nicolas},booktitle={Proceedings of the European Conference on Computer Vision (ECCV) 2nd Workshop on Marine Vision},year={2026},}
Promptable Animal Pose Tracking Across Species
Le Li, Daniela Ivanova, and Nicolas Pugeault
In Proceedings of the European Conference on Computer Vision (ECCV) Workshop on CV4Ecology, 2026
We present methods for tracking animal poses in videos using vision foundation models with limited labelled data. We propose two approaches: a supervised model employing a keypoint prompt encoder to explicitly inject structural priors from a reference frame, and an unsupervised approach leveraging foundation-model features for correspondence matching. The framework achieves strong results on the APTv2 and TigDog benchmarks, balancing accuracy with cross-species generalisation for wildlife conservation applications.
@inproceedings{le2026promptable,title={Promptable Animal Pose Tracking Across Species},author={Li, Le and Ivanova, Daniela and Pugeault, Nicolas},booktitle={Proceedings of the European Conference on Computer Vision (ECCV) Workshop on CV4Ecology},year={2026},}
Representation-Preserving Federated Self-Supervised Learning with Foundation Models
Özgü Göksu, Daniela Ivanova, and Nicolas Pugeault
In Proceedings of the European Conference on Computer Vision (ECCV) Workshop on LIMIT, 2026
@inproceedings{goksu2026federated,title={Representation-Preserving Federated Self-Supervised Learning with Foundation Models},author={Göksu, Özgü and Ivanova, Daniela and Pugeault, Nicolas},booktitle={Proceedings of the European Conference on Computer Vision (ECCV) Workshop on LIMIT},year={2026},note={Oral presentation},}
SceneHI: High-Resolution 3D-Consistent Scene Texturing with Controllable Illumination
Athanasios Tragakis, Marco Aversa, Daniela Ivanova, Chaitanya Kaul, Roderick Murray-Smith, Daniele Faccio, and Paul Henderson
In Proceedings of the European Conference on Computer Vision (ECCV), 2026
SceneHI is a framework that lifts high-resolution, illumination-aware priors from 2D diffusion models to perform 3D texture synthesis. It is the first to demonstrate that high-resolution textures, previously limited to 2D synthesis, can be generated directly on 3D objects without model fine-tuning or optimization. Designed for complex, multi-object environments, SceneHI uniquely combines 3D-consistency, high-resolution fidelity, and physically plausible baked shadows within a single generative pipeline. To enforce strict geometric coherence, we introduce an exact analytical pixel-to-texel mapping that aligns diffusion trajectories across multiple viewpoints. We utilize High-Resolution Latent Textures (HRLTs) as a persistent canvas for gradually denoised textures, while camera views perform the denoising steps in latent pixel space. This ensures a shared base texture that can be subsequently refined to high resolution without compromising multi-view consistency. Finally, a light-aware generative pass embeds realistic geometry-consistent shadows directly into the atlases, bridging the gap to production workflows. SceneHI achieves high visual fidelity while reducing generation time by 80% compared to existing scene-level methods.
@inproceedings{tragakis2026scenehi,title={SceneHI: High-Resolution 3D-Consistent Scene Texturing with Controllable Illumination},author={Tragakis, Athanasios and Aversa, Marco and Ivanova, Daniela and Kaul, Chaitanya and Murray-Smith, Roderick and Faccio, Daniele and Henderson, Paul},booktitle={Proceedings of the European Conference on Computer Vision (ECCV)},year={2026},}
Sampling 3D Gaussian Scenes in Seconds with Latent Diffusion Models
Paul Henderson, Melonie Almeida, Daniela Ivanova, and Titas Anciukevičius
In Proceedings of the International Joint Conference on Neural Networks (IJCNN), 2026
We present a latent diffusion model over 3D scenes, that can be trained using only 2D image data. To achieve this, we first design an autoencoder that maps multi-view images to 3D Gaussian splats, and simultaneously builds a compressed latent representation of these splats. Then, we train a multi-view diffusion model over the latent space to learn an efficient generative model. This pipeline does not require object masks nor depths, and is suitable for complex scenes with arbitrary camera positions. We conduct careful experiments on two large-scale datasets of complex real-world scenes – MVImgNet and RealEstate10K. We show that our approach enables generating 3D scenes in as little as 0.2 seconds, either from scratch, from a single input view, or from sparse input views. It produces diverse and high-quality results while running an order of magnitude faster than non-latent diffusion models and earlier NeRF-based generative models.
@inproceedings{henderson2026sampling,title={Sampling 3D Gaussian Scenes in Seconds with Latent Diffusion Models},author={Henderson, Paul and de Almeida, Melonie and Ivanova, Daniela and Anciukevičius, Titas},booktitle={Proceedings of the International Joint Conference on Neural Networks (IJCNN)},year={2026},}
Pixel-to-4D: Camera-Controlled Image-to-Video Generation with Dynamic 3D Gaussians
Existing camera-controlled image-to-video models often lack robust user controllability, and struggle with accurately modelling camera motion, maintaining temporal consistency, and preserving geometric integrity. We present a method for camera-controlled single-image to video generation with an intermediate 4D representation, achieving precise camera motion and temporal consistency. Our approach requires just one forward pass, without test-time optimisation or diffusion priors, and achieves state-of-the-art results on camera-controlled video prediction across four real-world datasets.
@inproceedings{dealmeida2026pixel,title={Pixel-to-4D: Camera-Controlled Image-to-Video Generation with Dynamic 3D Gaussians},author={de Almeida, Melonie and Ivanova, Daniela and Shi, Tong and Williamson, John and Henderson, Paul},booktitle={Proceedings of the International Conference on Pattern Recognition (ICPR)},year={2026},}
Splat-Portrait: Generalizing Talking Heads with Gaussian Splatting
Tong Shi, Melonie Almeida, Daniela Ivanova, Nicolas Pugeault, and Paul Henderson
In Proceedings of the International Conference on Multimedia Modeling (MMM), 2026
We present Splat-Portrait, a Gaussian-splatting-based method addressing the challenges of 3D head reconstruction and lip motion synthesis. Our approach automatically learns to disentangle a single portrait image into a static 3D reconstruction, represented as static Gaussian Splatting, and a predicted whole-image 2D background, then generates natural lip motion conditioned on input audio without any motion-driven priors. Training is driven purely by 2D reconstruction and score-distillation losses, without 3D supervision nor landmarks.
@inproceedings{shi2026splat,title={Splat-Portrait: Generalizing Talking Heads with Gaussian Splatting},author={Shi, Tong and de Almeida, Melonie and Ivanova, Daniela and Pugeault, Nicolas and Henderson, Paul},booktitle={Proceedings of the International Conference on Multimedia Modeling (MMM)},year={2026},}
Unsupervised Segmentation by Diffusing, Walking and Cutting
We propose an unsupervised image segmentation method using features from pre-trained text-to-image diffusion models. Inspired by classic spectral clustering approaches, we construct adjacency matrices from self-attention layers between image patches and recursively partition using Normalised Cuts. A key insight is that self-attention probability distributions, which capture semantic relations between patches, can be interpreted as a Markov random walk across the image. We leverage this by first using Random Walk Normalized Cuts directly on these self-attention activations to partition the image, minimizing transition probabilities across clusters while maximizing coherence within clusters. Applied recursively, this yields a hierarchical segmentation that reflects the rich semantics in the pre-trained attention layers, without any additional training. Next, we explore other ways to build the NCuts adjacency matrix from features, and how we can use the random walk interpretation of self-attention to capture long-range relationships. Finally, we propose an approach to automatically determine the NCut cost criterion, avoiding the need to tune this manually. We evaluate our method on a range of standard segmentation benchmarks, showing significant improvements over prior unsupervised approaches.
@inproceedings{ivanova2026unsupervised,title={Unsupervised Segmentation by Diffusing, Walking and Cutting},author={Ivanova, Daniela and Aversa, Marco and Henderson, Paul and Williamson, John},booktitle={Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)},year={2026},}
2025
ARTeFACT: Benchmarking Segmentation Models on Diverse Analogue Media Damage
Accurately detecting and classifying damage in analogue media such as paintings, photographs, textiles, mosaics, and frescoes is essential for cultural heritage preservation. While machine learning models excel in correcting global degradation if the damage operator is known a priori, we show that they fail to robustly predict where the damage is even after supervised training; thus, reliable damage detection remains a challenge. We introduce ARTeFACT, a dataset for damage detection in diverse types of analogue media, with over 11,000 annotations covering 15 kinds of damage across various subjects, media, and historical provenance. Furthermore, we contribute human-verified text prompts describing the semantic contents of the images, and derive additional textual descriptions of the annotated damage. We evaluate CNN, Transformer, diffusion-based segmentation models, and foundation vision models in zero-shot, supervised, unsupervised and text-guided settings, revealing their limitations in generalising across media types.
@inproceedings{ivanova2025artefact,title={ARTeFACT: Benchmarking Segmentation Models on Diverse Analogue Media Damage},author={Ivanova, Daniela and Aversa, Marco and Henderson, Paul and Williamson, John},booktitle={Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)},year={2025},}
2024
State-of-the-Art Fails in the Art of Damage Detection
Accurately detecting and classifying damage in analogue media such as paintings, photographs, textiles, mosaics, and frescoes is essential for cultural heritage preservation. While machine learning models excel in correcting global degradation if the damage operator is known a priori, we show that they fail to predict where the damage is even after supervised training; thus, reliable damage detection remains a challenge. We introduce DamBench, a dataset for damage detection in diverse analogue media, with over 11,000 annotations covering 15 damage types across various subjects, media, and historical provenance. We evaluate CNN, Transformer, and text-guided diffusion segmentation models, revealing their limitations in generalising across media types.
@inproceedings{ivanova2024sota,title={State-of-the-Art Fails in the Art of Damage Detection},author={Ivanova, Daniela and Aversa, Marco and Henderson, Paul and Williamson, John},booktitle={Proceedings of the European Conference on Computer Vision (ECCV) Workshop on VISART},year={2024},}
2023
Simulating analogue film damage to analyse and improve artefact restoration on high-resolution scans
Digital scans of analogue photographic film typically contain artefacts such as dust and scratches. Automated removal of these is an important part of preservation and dissemination of photographs of historical and cultural importance.
While state-of-the-art deep learning models have shown impressive results in general image inpainting and denoising, film artefact removal is an understudied problem. It has particularly challenging requirements, due to the complex nature of analogue damage, the high resolution of film scans, and potential ambiguities in the restoration. There are no publicly available high-quality datasets of real-world analogue film damage for training and evaluation, making quantitative studies impossible.
We address the lack of ground-truth data for evaluation by collecting a dataset of 4K damaged analogue film scans paired with manually-restored versions produced by a human expert, allowing quantitative evaluation of restoration performance. We construct a larger synthetic dataset of damaged images with paired clean versions using a statistical model of artefact shape and occurrence learnt from real, heavily-damaged images. We carefully validate the realism of the simulated damage via a human perceptual study, showing that even expert users find our synthetic damage indistinguishable from real. In addition, we demonstrate that training with our synthetically damaged dataset leads to improved artefact segmentation performance when compared to previously proposed synthetic analogue damage.
Finally, we use these datasets to train and analyse the performance of eight state-of-the-art image restoration methods on high-resolution scans. We compare both methods which directly perform the restoration task on scans with artefacts, and methods which require a damage mask to be provided for the inpainting of artefacts.
@article{ivanova23analogue,title={Simulating analogue film damage to analyse and improve artefact restoration on high-resolution scans},author={Ivanova, Daniela and Williamson, John and Henderson, Paul},journal={Computer Graphics Forum (Proceedings of Eurographics 2023)},year={2023},volume={42},number={2},doi={10.1111/cgf.14749},}
DiffInfinite: Large Mask-Image Synthesis via Parallel Random Patch Diffusion in Histopathology
Marco Aversa, Gabriel Nobis, Miriam Hägele, Kai Standvoss, Mihaela Chirica, Roderick Murray-Smith, Ahmed Alaa, Lukas Ruff, Daniela Ivanova, Wojciech Samek, and
3 more authors
Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks (NeurIPS), 2023
We present DiffInfinite, a hierarchical diffusion model that generates arbitrarily large histological images while preserving long-range correlation structural information. Our approach first generates synthetic segmentation masks, subsequently used as conditions for the high-fidelity generative diffusion process. The proposed sampling method can be scaled up to any desired image size while only requiring small patches for fast training. Moreover, it can be parallelized more efficiently than previous large-content generation methods while avoiding tiling artefacts. The training leverages classifier-free guidance to augment a small, sparsely annotated dataset with unlabelled data. Our method alleviates unique challenges in histopathological imaging practice: large-scale information, costly manual annotation, and protective data handling. The biological plausibility of DiffInfinite data is validated in a survey by ten experienced pathologists as well as a downstream segmentation task. Furthermore, the model scores strongly on anti-copying metrics which is beneficial for the protection of patient data.
@article{aversa2023diffinfinite,title={DiffInfinite: Large Mask-Image Synthesis via Parallel Random Patch Diffusion in Histopathology},author={Aversa, Marco and Nobis, Gabriel and Hägele, Miriam and Standvoss, Kai and Chirica, Mihaela and Murray-Smith, Roderick and Alaa, Ahmed and Ruff, Lukas and Ivanova, Daniela and Samek, Wojciech and Klauschen, Frederick and Sanguinetti, Bruno and Oala, Luis},journal={Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks (NeurIPS)},year={2023},doi={10.48550/arXiv.2306.13384},}
2022
Perceptual Loss based Approach for Analogue Film Restoration
Analogue film restoration, both for still photographs and motion picture emulsions,
is a slow and laborious manual process. Artifacts such as dust and scratches are random in shape,
size, and location; additionally, the ovserall degree of damage varies between different frames.
We address this less popular case of image restoration by training a U-Net model with a modified perceptual loss function.
Along with the novel perceptual loss function used for training, we propose a more rigorous quantitative model evaluation
approach which measures the overall degree of improvement in perceptual quality over our test set.
@article{ivanova2022,author={Ivanova, Daniela and Siebert, Jan and Williamson, John},title={Perceptual Loss based Approach for Analogue Film Restoration},journal={Proceedings of the 17th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications},year={2022},volume={4},pages={126-135},publisher={SciTePress},organization={INSTICC},doi={10.5220/0010829300003124},isbn={978-989-758-555-5},}