Reading Digits in Natural Images with Unsupervised Feature Learning
Explore this paper's citation graph
Summary
A new benchmark dataset for research use is introduced containing over 600,000 labeled digits cropped from Street View images, and variants of two recently proposed unsupervised feature learning methods are employed, finding that they are convincingly superior on benchmarks.
- Type
- article
- Published
- 2011-01-01
- Cited by
- 8,193
- References
- 29
- Access
- Open access
- OpenAlex
- https://openalex.org/W2335728318
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:16852518
Keywords
Computer science, Artificial intelligence, Feature (linguistics), Benchmark (surveying), Reading (process)
References
- Unified detection and recognition for reading text in scene images
- End-to-end text recognition with convolutional neural networks
- Google Street View: Capturing the World at Street Level
- The OCRopus open source OCR system
- An MRF Model for Binarization of Natural Scene Text
- End-to-end scene text recognition
- Learning hierarchical invariant spatio-temporal features for action recognition with independent subspace analysis
- Efficient implementation of local adaptive thresholding techniques using integral images
- Improvement of handwritten Japanese character recognition using weighted direction code histogram
- Text Detection and Character Recognition in Scene Images with Unsupervised Feature Learning
- Linear spatial pyramid matching using sparse coding for image classification
- Measuring Invariances in Deep Networks
- Unsupervised feature learning for audio classification using convolutional deep belief networks
- Large-scale privacy protection in Google Street View
- Deep, Big, Simple Neural Nets for Handwritten Digit Recognition
- Sparse deep belief net model for visual area V2
- A Robust System to Detect and Localize Texts in Natural Scene Images
- On-Line and Off-Line Handwriting Recognition: A Comprehensive Survey
- Historical review of OCR research and development
- A High-Throughput Screening Approach to Discovering Good Forms of Biologically Inspired Visual Representation
Cited by
- Regularization of Neural Networks using DropConnect
- Generalizing Pooling Functions in CNNs: Mixed, Gated, and Tree
- Enhance Visual Recognition Under Adverse Conditions via Deep Networks
- A Deep Learning Pipeline for Image Understanding and Acoustic Modeling
- Segmentation of brain MRI structures with deep machine learning
- Feature selection in computational biology
- Video Text Detection
- Path-SGD: Path-Normalized Optimization in Deep Neural Networks
- Deep learning of representations and its application to computer vision
- Scene text detection and recognition: recent advances and future trends
- Sequential prediction for budgeted learning : Application to trigger design. (Prédiction séquentielle pour l'apprentissage budgété : Application à la conception de trigger)
- A Deep Hashing Learning Network
- Deep extreme learning machines: supervised autoencoding architecture for classification
- Land Map Image Dataset: Ground-Truth And Classification Using Visual And Textural Features
- Faster SGD Using Sketched Conditioning
- An Analysis of the Connections Between Layers of Deep Neural Networks
- Text Recognition in Natural Images using Multiclass Hough Forests
- Training deep neural networks with low precision multiplications
- Principled Non-Linear Feature Selection
- End-to-end text recognition with convolutional neural networks