I am interested in multimodal representation learning, geospatial AI, and visual generation. My work focuses on applications in remote sensing, and I am also eager to explore a broader range of topics.
We propose Genesis, a generative engine that completes multi-scale satellite image pyramids from sparse seeds while preserving spatial and cross-scale consistency.
@inproceedings{khanal2026genesis,
title = {Genesis: A Generative Engine for Hierarchical Satellite Image
Synthesis},
author = {Khanal, Subash and Cui, Yangzhi and Cher, Daniel and Xing, Eric and Wei, Brian and Sastry, Srikumar and Jacobs, Nathan},
booktitle = {ACM SIGSPATIAL International Conference on Advances in Geographic Information Systems},
year = {2026}
}
We propose TerraDiT-Ω, a diffusion model for satellite imagery that is promptable with any geospatial primitive (polygons, polylines, bounding boxes, points) grounded in text.
@inproceedings{wei2026terraditomega,
title = {TerraDiT-$\Omega$: Unified Spatial Control for Satellite Image
Synthesis with Any Geospatial Primitive},
author = {Wei, Brian and Sastry, Srikumar and Cher, Daniel and Xing, Eric and Jacobs, Nathan},
booktitle = {European Conference on Computer Vision},
year = {2026}
}
We introduce TTE, a geolocation encoder that learns representations by tessellating the Earth's sphere using Voronoi partitioning.
@inproceedings{cher2026tte,
title = {Tesselating The Earth},
author = {Cher, Daniel and Iqbal, Hamza and Xing, Eric and Wei, Brian and Jacobs, Nathan},
booktitle = {European Conference on Computer Vision},
year = {2026}
}
We propose TerraDiT, a diffusion transformer for satellite image synthesis conditioned on sparse point locations and text — enabling annotation-efficient, spatially precise generation without dense pixel-level maps.
@article{sastry2026terradit,
title = {TerraDiT: Point-Conditioned Diffusion Transformer for Satellite Image Synthesis},
author = {Sastry, Srikumar and Cher, Daniel and Wei, Brian and Dhakal, Aayush and Khanal, Subash and Gupta, Dev and Jacobs, Nathan},
journal = {arXiv},
year = {2026}
}
We introduce VectorSynth, a diffusion-based model that generates pixel-accurate satellite imagery from polygonal geographic annotations with semantic attributes.
@inproceedings{cher2026vectorsynth,
title = {VectorSynth: Fine-Grained Satellite Image Synthesis with
Structured Semantics},
author = {Cher, Daniel and Wei, Brian and Sastry, Srikumar and Jacobs, Nathan},
booktitle = {Winter Conference on Applications of Computer Vision},
year = {2026}
}
I served as Co-President of the WashU Robotics Club, following my role as software lead on the rover project. Check out the club website for a look at the projects I was fortunate to work on during my time there!