THANK YOU FOR SUBSCRIBING
Recently, researchers have built machine learning and deep learning algorithms which can grasp the visual attributes of images, without the need for human involvement. It is imperative to train deep learning algorithms to achieve results in tasks pertaining to computer vision. Usually, the researchers need to exercise algorithms on datasets that are annotated in bulk, giving them versatile and detailed information of each image. Nevertheless, the time, resources and effort took to collect and annotate the images are significantly high.
The aim of the researchers involved in this program is to endow the computers the ability to read and comprehend text and textual data derived from any form of an image. Since humans use textual data to receive, send, understand and describe situations and information, it is important to provide machines with similar abilities, making it possible to reduce the amount of time and resources on large dataset annotations. The online context provides a supervisory signal, eliminating the need for specific annotations for individual images.
Stay ahead of the industry with exclusive feature stories on the top companies, expert insights and the latest news delivered straight to your inbox. Subscribe today.
Check out:Top Machine Learning Companies
In one of the studies, the researches prototyped computation models to collate data from online platforms that attach information of images with the visual data. Using these models, the algorithms are trained on the selection of viable attributes that illustrate images. Convolution neural networks (CNNs) form the base for another model where the separate layers are used to focus on different details such as that of pixel level details and abstract features. This technique provides an alternative to algorithms that are completely unsupervised by correlating images with elements that are not visual.
The researchers believe that using information available on the internet to train the algorithms will have its own significant advantages. By drawing useful information from a diverse pool of data, the amount of time taken to find new information has plunged.
Looking at the road ahead, the team aims to recognize better ways to use text-based information embedded in the image to answer queries about an image automatically. This can prove to be game changers in computer vision - saving time, cost, resources, and efforts.
Check out: Enterprise Technology Review
More in News