Surface consistency constraints in vision
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
Abstract Most computational theories of early visual processing require, as a first stage, the extraction of a symbolic representation (called theprimal sketch) of noticeable changes in image irradiance. For example, thezero-crossings of a Laplacian of a Gaussian operator applied to the image form the basis of one possible representation. Computational theories of stereo or motion correspondence only specify the computation of three-dimensional surface information at such points. Yet, the visual perception, consistent for different viewers, is clearly of complete surfaces. Since in principle the class of surface which could pass through the known boundary points provided by feature point correspondence is infinite and contains widely varying surfaces, the visual system must incorporate some additional constraints besides the known points to compute the complete surface. Using the image irradiance equation, asurface consistency constraint, referred to informally asno news is good news is derived. The constraint implies that the surface must agree with the correspondence information, andnot vary radically between these points. An explicit form of this surface consistency constraint is derived, by relating the probability of a zero-crossing in a region of the image to the variation in the local surface orientation of the surface, provided that the surface albedo and the illumination are roughly constant. The surface consistency constraint is informally supported by Logan's theorem, which essentially states that all the critical information of a signal is generally contained in its zero-crossings, and by the demonstration that the transformation from surface shape to image irradiance generally preserves zero-crossings. Hence, it is argued that surface shape is implicitly encoded in the positions of the zero-crossings of the convolved images, and further, that the encoding can be at least partially inverted to construct an approximation to the viewed surface from the correspondence information.
