“I don’t think I have ever done anything as peculiar in my life. Among other things, it shows a young man looking with interest at a print on the wall of an exhibition that features himself. How can this be? Perhaps I am not far removed from Einstein’s curved universe.” So wrote M. C. Escher about his 1956 lithograph Print Gallery. Nearly half a century later, a mathematical analysis related its geometry to an untwisted source image through a conformal power map $z \mapsto z^{\alpha},\; \alpha \in \mathbb{C}$. Building on this construction, we use a frozen text-to-image diffusion model to generate new self-referential scenes. Prompting alone does not enforce the recursion, while a post-hoc transformation can leave structures poorly connected. Applying the transformation during sampling is also insufficient: the denoiser may “repair” the intended distortion or drift out of the prescribed geometry. We construct a generalized inverse $T^{\dagger}$ of the non-invertible image transformation $T$, adapted to its recursive constraint. In the idealized formulation, the Penrose identity $TT^{\dagger}T = T$ makes $TT^{\dagger}$ an idempotent projection onto geometrically admissible images. Yet denoising only the transformed image remains an out-of-distribution task, even with projection. We therefore braid denoising steps with $T$ and $T^{\dagger}$: source-space steps develop the untwisted scene, while transformed-space steps refine its appearance and connections in the final geometry. We generate Print Gallery-like compositions and explore further transformations. Rather than distorting a finished image, we let the scene and its distortion develop together.
Assaf Shocher
I am an Assistant Professor at the Technion in the Faculty of Data and Decision Sciences. Previously, I was a Postdoctoral Research Scientist at NVIDIA, a postdoctoral researcher at UC Berkeley with Alyosha Efros, and a Visiting Scholar at Google DeepMind. I received my PhD from the Weizmann Institute of Science, advised by Michal Irani, and hold bachelor’s degrees in Physics and Electrical Engineering from Ben-Gurion University. More details in About.
My research focuses on computer vision and deep learning. I aim to bridge theory and practical application in machine learning. While admiring engineering advances, I am drawn to the scientific investigation of foundational principles. Fascinated by elegant ideas and mathematical observations, I start each project from first principles to develop methods that offer fundamentally new perspectives on problems. In particular, I study algebraic properties of neural networks, including analogues of inverses and projections, to make them easier to analyze, compose, and control, with applications to inverse problems and adaptive learning.
Selected publications
19 publications
The Moore-Penrose Pseudo-inverse (PInv) serves as the fundamental solution for linear systems. In this paper, we propose a natural generalization of PInv to the nonlinear regime in general and to neural networks in particular. We introduce Surjective Pseudo-invertible Neural Networks (SPNN), a class of architectures explicitly designed to admit a tractable non-linear PInv. The proposed non-linear PInv and its implementation in SPNN satisfy fundamental geometric properties. One such property is null-space projection or "Back-Projection", x′=x+A†(y−Ax), which moves a sample x to its closest consistent state x′ satisfying Ax=y. We formalize Non-Linear Back-Projection (NLBP), a method that guarantees the same consistency constraint for non-linear mappings f(x)=y via our defined PInv. We leverage SPNNs to expand the scope of zero-shot inverse problems. Diffusion-based null-space projection has revolutionized zero-shot solving for linear inverse problems by exploiting closed-form back-projection. We extend this method to non-linear degradations. Here, "degradation" is broadly generalized to include any non-linear loss of information, spanning from optical distortions to semantic abstractions like classification. This approach enables zero-shot inversion of complex degradations and allows precise semantic control over generative outputs without retraining the diffusion prior.
Neural networks are famously nonlinear. However, linearity is defined relative to a pair of vector spaces, $f:\mathcal{X}\rightarrow\mathcal{Y}.$ Is it possible to identify a pair of non-standard vector spaces for which a conventionally nonlinear function is, in fact, linear? This paper introduces a method that makes such vector spaces explicit by construction. We find that if we sandwich a linear operator A between two invertible neural networks, $f(x)=g_{y}^{-1}(Ag_{x}(x))$, then the corresponding vector spaces X and Y are induced by newly defined addition and scaling actions. This framework makes the entire arsenal of linear algebra, including SVD, pseudo-inverse, and more, applicable to nonlinear mappings. We demonstrate this by collapsing diffusion model sampling into a single step, enforcing global idempotency, and demonstrating modular style transfer.
Deep learning models often struggle when deployed in real-world settings due to distribution shifts between training and test data. We present Idempotent Test-Time Training ($IT^{3}$), a novel approach that enables on-the-fly adaptation to distribution shifts using only the current test instance, without any auxiliary task design. Our key insight is that enforcing idempotence—where repeated applications of a function yield the same result—can effectively replace domain-specific auxiliary tasks used in previous TTT methods. We theoretically connect idempotence to prediction confidence and demonstrate that minimizing the distance between successive applications of our model during inference leads to improved out-of-distribution performance. Our results suggest that idempotence provides a universal principle for test-time adaptation that generalizes across domains and architectures.
Traditional super-resolution (SR) methods assume an "ideal" downscaling SR-kernel (e.g., bicubic). Such methods fail once the LR images are generated differently. Current blind-SR methods are still fundamentally restricted to rather simplistic downscaling SR-kernels. In "KernelFusion" we introduce a zero-shot diffusion-based method that makes no assumptions about the kernel. Our method recovers the unique image-specific SR-kernel directly from the LR input image, while simultaneously recovering its corresponding HR image. KernelFusion exploits the principle that the correct SR-kernel is the one that maximizes patch similarity across different scales of the LR image. By breaking free from predefined kernel assumptions, KernelFusion pushes Blind-SR into a new assumption-free paradigm, handling downscaling kernels previously thought impossible.
Video encoders optimize compression for human perception. In many modern applications, videos serve as input for AI systems performing tasks like object recognition. It is therefore useful to optimize the encoder for a downstream task. A major challenge is how to combine such optimization with existing standard video encoders. Here, we address this challenge by controlling the Quantization Parameters (QPs) at the macro-block level to optimize the downstream task. This granular control allows us to prioritize encoding for task-relevant regions. We formulate this as a Reinforcement Learning (RL) task, where the agent learns to balance long-term implications of choosing QPs on both task performance and bit-rate constraints.
We propose a new approach for generative modeling based on training a neural network to be idempotent. An idempotent operator is one that can be applied sequentially without changing the result beyond the initial application, namely $f(f(z))=f(z)$. The proposed model $f$ is trained to map a source distribution (e.g, Gaussian noise) to a target distribution (e.g. realistic images) using the following objectives: (1) Instances from the target distribution should map to themselves, namely $f(x)=x$. (2) Instances from the source distribution should map onto the defined target manifold. This is achieved by optimizing the idempotence term, $f(f(z))=f(z)$. This strategy results in a model capable of generating an output in one step, maintaining a consistent latent space, while also allowing sequential applications for refinement. This work is a first step towards a "global projector" that enables projecting any input into a target data distribution.
Text-to-image diffusion models have demonstrated an unparalleled ability to generate high-quality, diverse images from a textual prompt. However, the internal representations learned by these models remain an enigma. In this work, we present Conceptor, a novel method to interpret the internal representation of a textual concept by a diffusion model. This interpretation is obtained by decomposing the concept into a small set of human-interpretable textual elements. Applied over the state-of-the-art Stable Diffusion model, Conceptor reveals non-trivial structures in the representations of concepts, such as surprising visual connections between concepts, biases, renowned artistic styles, or a simultaneous fusion of multiple meanings of the concept.
Improving correlation based super-resolution microscopy images through image fusion by self-supervised deep learning
Optics Express 2024
Super-resolution imaging is a powerful tool in modern biological research, allowing for the optical observation of subcellular structures with great detail. In this paper, we present a deep learning approach for image fusion of intensity and super-resolution optical fluctuation imaging (SOFI) microscopy images. We construct a network that can successfully combine the advantages of these two imaging methods, producing a fused image with a resolution comparable to that of SOFI and an SNR comparable to that of the intensity image. We also demonstrate the effectiveness of our approach experimentally. Our network is designed as a self-supervised network and shows the ability to train on a single pair of images and to generalize to other image pairs without the need for additional training. Our approach offers a flexible and efficient way to combine the strengths of correlation based imaging techniques along with traditional intensity based microscopy.
Masked Image Modeling (MIM) is a promising self-supervised learning approach that enables learning from unlabeled images. Despite its recent success, learning good representations through MIM remains challenging because it requires predicting the right semantic content in accurate locations. For example, given an incomplete picture of a dog, we can guess that there is a tail, but we cannot determine its exact location. In this work, we propose to incorporate location uncertainty into MIM by using stochastic positional embeddings (StoP). Specifically, we condition the model on stochastic masked token positions drawn from a Gaussian distribution. StoP reduces overfitting to location features and guides the model toward learning features that are more robust to location uncertainties. Quantitatively, StoP improves downstream MIM performance on a variety of downstream tasks, including +1.7% on ImageNet linear probing using ViT-B, and +2.5% for ViT-H using 1% of the data.
Do different neural networks, trained for various vision tasks, share some common representations? In this paper, we demonstrate the existence of common features we call "Rosetta Neurons" across a range of models with different architectures, different tasks (generative and discriminative), and different types of supervision. Our findings suggest that certain visual concepts and structures are inherently embedded in the natural world and can be learned by different models regardless of the specific task or architecture. The Rosetta Neurons facilitate model-to-model translation enabling various inversion-based manipulations, including cross-class alignments, shifting, zooming, and more, without the need for specialized training.
Single video GANs require unreasonable amount of time to train, rendering them almost impractical. In this paper we question the necessity of a GAN for generation from a single video, and introduce a non-parametric baseline for a variety of generation and manipulation tasks. We revive classical space-time patches-nearest-neighbors approaches and adapt them to a scalable unconditional generative model, without any learning. This simple baseline surprisingly outperforms single-video GANs in visual quality and realism, and is disproportionately faster (runtime reduced from several days to seconds). These observations show that the classical approaches, if adapted correctly, significantly outperform heavy deep learning machinery for these tasks.
Drop the GAN: In Defense of Patches Nearest Neighbors as Single Image Generative Models
CVPR 2022 (Oral)
Single image generative models perform synthesis and manipulation tasks by capturing the distribution of patches within a single image. The classical prevailing approaches are based on an optimization process that maximizes patch similarity. Recently, Single Image GANs were introduced as a superior solution. In this paper, we show that all of these tasks can be performed without any training, within several seconds, in a unified, surprisingly simple framework. We revisit and cast the "good-old" patch-based methods into a novel optimization-free framework. Not only is our method faster (×10³-×10⁴ than a GAN), it produces superior results, less artifacts and more realistic global structure than any of the previous approaches.
A basic operation in Convolutional Neural Networks (CNNs) is spatial resizing of feature maps. This is done either by strided convolution (downscaling) or transposed convolution (upscaling). Such operations are limited to a fixed filter moving at predetermined integer steps. We propose a generalization of the common Conv-layer, from a discrete layer to a Continuous Convolution (CC) Layer. CC Layers naturally extend Conv-layers by representing the filter as a learned continuous function over sub-pixel coordinates. This allows learnable and principled resizing of feature maps, to any size, dynamically and consistently across scales. Once trained, the CC layer can be used to output any scale/size chosen at inference time.
We present a novel GAN-based model that utilizes the space of deep features learned by a pre-trained classification model. Inspired by classical image pyramid representations, we construct our model as a Semantic Generation Pyramid - a hierarchical framework which leverages the continuum of semantic information encapsulated in such deep features. Given a set of features extracted from a reference image, our model generates diverse image samples, each with matching features at each semantic level of the classification model. We demonstrate that our model results in a versatile and flexible framework that can be used in various classic and novel image generation tasks.
Super-resolution (SR) methods typically assume that the low-resolution (LR) image was downscaled from the unknown high-resolution (HR) image by a fixed 'ideal' downscaling kernel. In this paper we introduce "KernelGAN", an image-specific Internal-GAN, which trains solely on the LR test image at test time, and learns its internal distribution of patches. Its Generator is trained to produce a downscaled version of the LR test image, such that its Discriminator cannot distinguish between the patch distribution of the downscaled image, and the patch distribution of the original LR image. The Generator, once trained, constitutes the downscaling operation with the correct image-specific SR-kernel.
Generative Adversarial Networks (GANs) typically learn a distribution of images in a large image dataset. In this paper we propose an "Internal GAN" (InGAN) - an image-specific GAN - which trains on a single input image and learns its internal distribution of patches. It is then able to synthesize a plethora of new natural images of significantly different sizes, shapes and aspect-ratios - all with the same internal patch-distribution (same "DNA") as the input image. InGAN is fully unsupervised, requiring no additional data other than the input image itself.
Many seemingly unrelated computer vision tasks can be viewed as a special case of image decomposition into separate layers. In this paper we propose a unified framework for unsupervised layer decomposition of a single image, based on coupled "Deep-image-Prior" (DIP) networks. We show that coupling multiple such DIPs provides a powerful tool for decomposing images into their basic components, for a wide variety of applications. This capability stems from the fact that the internal statistics of a mixture of layers is more complex than the statistics of each of its individual components. We show the power of this approach for Image-Dehazing, Fg/Bg Segmentation, Watermark-Removal, and more, in a totally unsupervised way.
Deep Learning has led to a dramatic leap in SuperResolution (SR) performance. However, being supervised, these methods are restricted to specific training data. Real LR images, however, rarely obey these restrictions, resulting in poor SR results. In this paper we introduce "Zero-Shot" SR, which exploits the power of Deep Learning, but does not rely on prior training. We exploit the internal recurrence of information inside a single image, and train a small image-specific CNN at test time, on examples extracted solely from the input image itself. As such, it can adapt itself to different settings per image. This allows to perform SR of real old photos, noisy images, biological data, and other images where the acquisition process is unknown or non-ideal.