Sound Retrieval and Ranking Using Sparse Auditory Representations
Explore this paper's citation graph
Summary
A machine-vision method is adapted, the passive-aggressive model for image retrieval (PAMIR), which efficiently learns a linear mapping from a very large sparse feature space to a large query-term space and shows a significant advantage for the auditory models over vector-quantized MFCCs.
- Type
- article
- Published
- 2010-09-01
- Cited by
- 65
- References
- 32
- OpenAlex
- https://openalex.org/W1966273763
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:8542932
Keywords
Computer science, Mel-frequency cepstrum, Computational auditory scene analysis, Speech recognition, Ranking (information retrieval)
References
- On the importance of time—a temporal representation of sound
- A computational model of filtering, detection, and compression in the cochlea
- Large-scale kernel machines
- Vector quantization and signal compression
- The Mammalian Auditory Pathway: Neurophysiology
- Comparison of the roex and gammachirp filters as representations of the auditory filter
- Multi-Tasking with Joint Semantic Spaces for Large-Scale Music Annotation and Retrieval
- Audio Information Retrieval using Semantic Similarity
- Auditory images:How complex sounds are represented in the auditory system
- A human nonlinear cochlear filterbank.
- Large-scale content-based audio retrieval from text queries
- The lower limit of pitch as determined by rate discrimination.
- Derivation of auditory filter shapes from notched-noise data.
- Recursive nearest neighbor search in a sparse and multiscale domain for comparing audio signals
- Sparse coding of sensory inputs.
- Cascades of two-pole-two-zero asymmetric resonators are good models of peripheral auditory function.
- An analog electronic cochlea
- Semantic Annotation and Retrieval of Music and Sound Effects
- Emergence of simple-cell receptive field properties by learning a sparse code for natural images
- Musical query-by-description as a multiclass learning problem
Cited by
- Discriminative latent variable models for visual recognition
- Reconnaissance des sons de l'environnement dans un contexte domotique. (Environmental sounds recognition in a domotic context)
- A real-time implementation of the primary auditory neuron activities
- Recognizing and Classifying Environmental Sounds
- Statistical distribution of common audio features: encounters in a heavy-tailed universe
- Large-Scale Music Annotation and Retrieval: Learning to Rank in Joint Semantic Spaces
- Tone-vs.-Noise Dichotomy in Psychoacoustics Could be Resolved by Careful Consideration of Cochlear Signal Processing
- Discriminative tag learning on YouTube videos with latent sub-tags
- Robust Sound Event Classification Using Deep Neural Networks
- A Bag-of-Features Framework to Classify Time Series
- Multi-Tasking with Joint Semantic Spaces for Large-Scale Music Annotation and Retrieval
- A Pole-Zero Filter Cascade Provides Good Fits to Human Masking Data and to Basilar Membrane and Neural Data
- [self.]: an Interactive Art Installation that Embodies Artificial Intelligence and Creativity
- An Overview on Perceptually Motivated Audio Indexing and Classification
- A Systematic Evaluation of the Bag-of-Frames Representation for Music Information Retrieval
- Codebook-Based Audio Feature Representation for Music Information Retrieval
- Automatically Discovering Talented Musicians with Acoustic Analysis of YouTube Videos
- Feature learning and deep architectures: new directions for music informatics
- [self.]: an Interactive Art Installation that Embodies Artificial Intelligence and Creativity: A Demonstration
- Acoustic scene classification using sparse feature learning and event-based pooling
Related papers
- The Research Advance of Auditory Scene Analysis
- Cluster Analysis for the Separation of Auditory Scenes
- Modeling binaural auditory scene analysis by a temporal fuzzy cluster analysis approach
- A really complicated problem: Auditory scene analysis
- Modelling auditory scene analysis: a representational approach
- Auditory blobs
- The auditory organization of speech and other sources in listeners and computational models