With the rise in popularity of 3D entertainment, demand for 3D content is growing. Current options to generate 3D content are often really pricey because they either require very advanced equipment to film in 3D or a considerable amount of time to convert 2D content to 3D content by adding information about the depth to the image. The goal of this research is to find a semi-automatic technique to convert 2D content to 3D content that is faster and that generates satisfactory results to use the depth to convert to 3D. To attain this goal, a technique was developed to allow the user to only annotate quickly part of a key frame in a video sequence and then propagate this information to estimate the depth of all pixels in the first image and then through the similar images in the video sequence. Effectively allowing the user to only annotate a few images in a video to convert the whole video to 3D.
The technique uses primarly the Random Walker method to propagate the information from the annotations to the entire image. This phase can be compared to a multi-class graph segmentation problem. Then by developing an iterative method, we can separate the second phase which consists of the propagation of the information of depth that was computed on the first image to the rest of the images in the sequence. The proposed approach considers the image as a graph where the nodes corresponds to the pixels and the edges have a computed weight that corresponds to the similarity of the pixels that are linked. Results obtained with this method show that it performs well on a single image with only a small number of annotations from the user. However, the propagation of the information to other images faces multiple issues. For example, problems can come from the occlusions in an image, the errors are also propagated at the same time as the information in the image and finally there can be problems with variations in illumination in the image.
Following these problems, the research was oriented towards finding a good method to compute the motion between two images to find the exact stereo-correspondence by motion estimation. The results of this research was an automatic dense motion estimation method, which means finding the displacement of each pixels between two images. The proposed approach uses the similarity between neighboring pixels in the same image and the similarity between pixels included in a research kernel in the following image in the sequence to determine the most likely motion for the current pixel. Found motions are kept as a probability and pixels are influenced by other neighboring similar pixels to "agree" on a shared motion. Similarity between pixels is computed using each components of the Lab color space of the current pixels and also a small number of neighboring pixels and the neighboring gradients. Results obtained with the method are compared against a published dataset from Middleburry that regroups motion estimation methods. Test images are determined and two measures of success are defined to compare methods. The comparison with other methods achieves good results on the test images but the visual results shows some small regions containing errors. The performance of the method was further improved by considering SIFT stereo-correspondence, matching results and the speed of the method was improved by regrouping similar pixels under a region called a superpixel. Only a small number of pixels in this superpixel are considered to reduce the total number of computations.
Finally the different effects of the control parameters of the method are explained in another section. And the difference in results with and without SIFT, or with or without the superpixels are also explained. The main problems of the proposed method are explained and some solutions are proposed even if they were not implemented.
Rocheleau, É. (Author),
Desrosiers (Supervisor) &
Vázquez (Co-supervisor),
21 Jun 2017Student thesis: Master's thesis › Master in Engineering: Engineering