248 papers
eess as
0/02026
3- JulWaveNet-Style Guitar Amplifier Model Pruning for Real-Time iOS Deploymentno summary yeteess-as2607.10086Stanford0 citesJul 11, 2026
- JunFlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speechno summary yeteess-as2606.23190Alibaba0 citesJun 22, 2026
- JunPreserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generationno summary yeteess-as2606.30944Microsoft Research0 citesJun 29, 2026
2024
22023
10- DecIR-UWB Radar-Based Contactless Silent Speech Recognition of Vowels, Consonants, Words, and Phrasesno summary yeteess-as2312.09572NVIDIA0 citesDec 15, 2023
- SepPESTO: Pitch Estimation with Self-supervised Transposition-equivariant Objectiveno summary yeteess-as2309.02265Sony2 citesSep 5, 2023
- SepA Generalized Bandsplit Neural Network for Cinematic Audio Source Separationno summary yeteess-as2309.02539Netflix1 citesSep 5, 2023
- AugThe Sound Demixing Challenge 2023 $\unicode{x2013}$ Music Demixing Trackno summary yeteess-as2308.06979Tencent0 citesAug 14, 2023
- AugUltra Dual-Path Compression For Joint Echo Cancellation And Noise Suppressionno summary yeteess-as2308.11053Tencent5 citesAug 21, 2023
- JulGlobal birdsong embeddings enable superior transfer learning for bioacoustic classificationno summary yeteess-as2307.06292Google Research4 citesJul 12, 2023
- JunLearning When to Trust Which Teacher for Weakly Supervised ASRno summary yeteess-as2306.12012Amazon0 citesJun 21, 2023
- JunConfidence-based Ensembles of End-to-End Speech Recognition Modelsno summary yeteess-as2306.15824NVIDIA0 citesJun 27, 2023
- MarTS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker Embeddingsno summary yeteess-as2303.03849Microsoft Research2 citesMar 7, 2023
- MarClinical BERTScore: An Improved Measure of Automatic Speech Recognition Performance in Clinical Settingsno summary yeteess-as2303.05737Google Research1 citesMar 10, 2023
2022
6- AprMusic Source Separation with Generative Flowno summary yeteess-as2204.09079Tencent8 citesApr 19, 2022
- MarHarmonicity Plays a Critical Role in DNN Based Versus in Biologically-Inspired Monaural Speech Segregation Systemsno summary yeteess-as2203.04420Google Research1 citesMar 8, 2022
- FebRescoreBERT: Discriminative Speech Recognition Rescoring with BERTno summary yeteess-as2202.01094Amazon36 citesFeb 2, 2022
- FebContrastive-mixup learning for improved speaker verificationno summary yeteess-as2202.10672Amazon9 citesFeb 22, 2022
- FebImproving fairness in speaker verification via Group-adapted Fusion Networkno summary yeteess-as2202.11323Amazon12 citesFeb 23, 2022
- FebopenFEAT: Improving Speaker Identification by Open-set Few-shot Embedding Adaptation with Transformerno summary yeteess-as2202.12349Amazon10 citesFeb 24, 2022
2021
51- NovUnsupervised Speech Enhancement with speech recognition embedding and disentanglement lossesno summary yeteess-as2111.08678Microsoft Research1 citesNov 16, 2021
- NovA Conformer-based ASR Frontend for Joint Acoustic Echo Cancellation, Speech Enhancement and Speech Separationno summary yeteess-as2111.09935Google Research1 citesNov 18, 2021
- NovRepresentation learning through cross-modal conditional teacher-student training for speech emotion recognitionno summary yeteess-as2112.00158Amazon0 citesNov 30, 2021
- OctCTC Variations Through New WFST Topologiesno summary yeteess-as2110.03098NVIDIA14 citesOct 6, 2021
- OctSimilarity-and-Independence-Aware Beamformer with Iterative Casting and Boost Start for Target Source Extraction Using Referenceno summary yeteess-as2110.09019Sony5 citesOct 18, 2021
- SepImproving Speaker Identification for Shared Devices by Adapting Embeddings to Speaker Subsetsno summary yeteess-as2109.02576Amazon5 citesSep 6, 2021
- SepRemember the context! ASR slot error correction through memorizationno summary yeteess-as2109.05092Amazon1 citesSep 10, 2021
- SepBigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognitionno summary yeteess-as2109.13226Google Research158 citesSep 27, 2021
- SepThe impact of non-target events in synthetic soundscapes for sound event detectionno summary yeteess-as2109.14061Google Research3 citesSep 28, 2021
- AugAmortized Neural Networks for Low-Latency Speech Recognitionno summary yeteess-as2108.01553Amazon0 citesAug 3, 2021
- AugExploring Retraining-Free Speech Recognition for Intra-sentential Code-Switchingno summary yeteess-as2109.00921Apple5 citesAug 27, 2021
- JulMulti-user VoiceFilter-Lite via Attentive Speaker Embeddingno summary yeteess-as2107.01201Google Research5 citesJul 2, 2021
- JulComparing Supervised Models And Learned Speech Representations For Classifying Intelligibility Of Disordered Speech On Selected Phrasesno summary yeteess-as2107.03985Google Research0 citesJul 8, 2021
- JulLow complexity online convolutional beamformingno summary yeteess-as2107.06775Microsoft Research0 citesJul 14, 2021
- JunA Neural Acoustic Echo Canceller Optimized Using An Automatic Speech Recognizer And Large Scale Synthetic Datano summary yeteess-as2106.00856Google Research0 citesJun 1, 2021
- JunTeaching keyword spotters to spot new keywords with limited examplesno summary yeteess-as2106.02443Google Research2 citesJun 4, 2021
- JunHuman Listening and Live Captioning: Multi-Task Training for Speech Enhancementno summary yeteess-as2106.02896Microsoft Research2 citesJun 5, 2021
- JunWeakly-supervised word-level pronunciation error detection in non-native English speechno summary yeteess-as2106.03494Amazon0 citesJun 7, 2021
- JunPersonalized PercepNet: Real-time, Low-complexity Target Voice Separation and Enhancementno summary yeteess-as2106.04129Amazon4 citesJun 8, 2021
- JunMulti-channel Opus compression for far-field automatic speech recognition with a fixed bitrate budgetno summary yeteess-as2106.07994Amazon0 citesJun 15, 2021
- JunScaling Laws for Acoustic Modelsno summary yeteess-as2106.09488Amazon1 citesJun 11, 2021
- JunASR Adaptation for E-commerce Chatbots using Cross-Utterance Context and Multi-Task Language Modelingno summary yeteess-as2106.09532Amazon4 citesJun 15, 2021
- JunOn-Device Personalization of Automatic Speech Recognition Models for Disordered Speechno summary yeteess-as2106.10259Google Research12 citesJun 18, 2021
- MayPoint Cloud Audio Processingno summary yeteess-as2105.02469Adobe0 citesMay 6, 2021
- MayDifferentiable Signal Processing With Black-Box Audio Effectsno summary yeteess-as2105.04752Adobe7 citesMay 11, 2021
- MayListen with Intent: Improving Speech Recognition with Audio-to-Intent Front-Endno summary yeteess-as2105.07071Amazon2 citesMay 14, 2021
- MayDPLM: A Deep Perceptual Spatial-Audio Localization Metricno summary yeteess-as2105.14180Meta / FAIR1 citesMay 29, 2021
- AprSpeakerStew: Scaling to Many Languages with a Triaged Multilingual Text-Dependent and Text-Independent Speaker Verification Systemno summary yeteess-as2104.02125Google Research2 citesApr 5, 2021
- AprPushing the Limits of Non-Autoregressive Speech Recognitionno summary yeteess-as2104.03416Google Research17 citesApr 7, 2021
- AprMulti-Task Learning for End-to-End ASR Word and Utterance Confidence with Deletion Predictionno summary yeteess-as2104.12870Google Research2 citesApr 26, 2021
- MarContrastive Separative Coding for Self-supervised Representation Learningno summary yeteess-as2103.00816Tencent0 citesMar 1, 2021
- MarSandglasset: A Light Multi-Granularity Self-attentive Network For Time-Domain Speech Separationno summary yeteess-as2103.00819Tencent4 citesMar 1, 2021
- MarBest of Both Worlds: Robust Accented Speech Recognition with Adversarial Transfer Learningno summary yeteess-as2103.05834Amazon0 citesMar 10, 2021
- MarLearning Word-Level Confidence For Subword End-to-End ASRno summary yeteess-as2103.06716Google Research2 citesMar 11, 2021
- MarWav2vec-C: A Self-supervised Model for Speech Representation Learningno summary yeteess-as2103.08393Amazon3 citesMar 9, 2021
- MarTarget Speaker Verification with Selective Auditory Attention for Single and Multi-talker Speechno summary yeteess-as2103.16269Tencent2 citesMar 30, 2021
- FebInternal Language Model Training for Domain-Adaptive End-to-End Speech Recognitionno summary yeteess-as2102.01380Microsoft Research1 citesFeb 2, 2021
- FebEnd-to-End Multi-Channel Transformer for Speech Recognitionno summary yeteess-as2102.03951Amazon1 citesFeb 8, 2021
- FebCDPAM: Contrastive learning for perceptual audio similarityno summary yeteess-as2102.05109Adobe3 citesFeb 9, 2021
- FebLow-Complexity, Real-Time Joint Neural Echo Control and Speech Enhancement Based On PercepNetno summary yeteess-as2102.05245Amazon2 citesFeb 10, 2021
- FebEnhancing into the codec: Noise Robust Speech Coding with Vector-Quantized Autoencodersno summary yeteess-as2102.06610Amazon2 citesFeb 12, 2021
- FebOn training targets for noise-robust voice activity detectionno summary yeteess-as2102.07445Microsoft Research2 citesFeb 15, 2021
- FebSemi-Supervised Singing Voice Separation with Noisy Self-Trainingno summary yeteess-as2102.07961Amazon1 citesFeb 16, 2021
- FebContext-Aware Prosody Correction for Text-Based Speech Editingno summary yeteess-as2102.08328Adobe3 citesFeb 16, 2021
- FebGenerative Speech Coding with Predictive Variance Regularizationno summary yeteess-as2102.09660Google Research8 citesFeb 18, 2021
- FebWARP-Q: Quality Prediction For Generative Neural Speech Codecsno summary yeteess-as2102.10449Google Research2 citesFeb 20, 2021
- FebHandling Background Noise in Neural Speech Generationno summary yeteess-as2102.11906Google Research0 citesFeb 23, 2021
- JanEffective Low-Cost Time-Domain Audio Separation Using Globally Attentive Locally Recurrent Networksno summary yeteess-as2101.05014Tencent9 citesJan 13, 2021
- JanTiny Transducer: A Highly-efficient Speech Recognition Model on Edge Devicesno summary yeteess-as2101.06856Tencent2 citesJan 18, 2021
- JanTowards efficient models for real-time deep noise suppressionno summary yeteess-as2101.09249Microsoft Research8 citesJan 22, 2021
- JanA Review of Speaker Diarization: Recent Advances with Deep Learningno summary yeteess-as2101.09624Microsoft Research41 citesJan 24, 2021
2020
111- DecThe Third DIHARD Diarization Challengeno summary yeteess-as2012.01477Baidu22 citesDec 2, 2020
- DecREDAT: Accent-Invariant Representation for End-to-End ASR by Domain Adversarial Training with Relabelingno summary yeteess-as2012.07353Amazon2 citesDec 14, 2020
- DecDenoiSpeech: Denoising Text to Speech with Frame-Level Noise Modelingno summary yeteess-as2012.09547Microsoft Research1 citesDec 17, 2020
- DecMulti-channel Multi-frame ADL-MVDR for Target Speech Separationno summary yeteess-as2012.13442Tencent5 citesDec 24, 2020
- DecDetection of Lexical Stress Errors in Non-Native (L2) English with Data Augmentation and Attentionno summary yeteess-as2012.14788Amazon1 citesDec 29, 2020
- NovDOVER-Lap: A Method for Combining Overlap-aware Diarization Outputsno summary yeteess-as2011.01997Amazon9 citesNov 3, 2020
- NovIntegration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysisno summary yeteess-as2011.02014Google Research1 citesNov 3, 2020
- NovOne-shot conditional audio filtering of arbitrary soundsno summary yeteess-as2011.02421Google Research2 citesNov 4, 2020
- NovMinimum Bayes Risk Training for End-to-End Speaker-Attributed ASRno summary yeteess-as2011.02921Microsoft Research2 citesNov 3, 2020
- NovExploring End-to-End Multi-channel ASR with Bias Information for Meeting Transcriptionno summary yeteess-as2011.03110Microsoft Research0 citesNov 5, 2020
- NovListen, Look and Deliberate: Visual context-aware speech recognition using pre-trained text-video representationsno summary yeteess-as2011.04084Microsoft Research2 citesNov 8, 2020
- NovBenchmarking LF-MMI, CTC and RNN-T Criteria for Streaming ASRno summary yeteess-as2011.04785Meta / FAIR0 citesNov 9, 2020
- NovFastSVC: Fast Cross-Domain Singing Voice Conversion with Feature-wise Linear Modulationno summary yeteess-as2011.05731Tencent2 citesNov 11, 2020
- NovEfficient Knowledge Distillation for RNN-Transducer Modelsno summary yeteess-as2011.06110Google Research3 citesNov 11, 2020
- NovAudio-visual Multi-channel Integration and Recognition of Overlapped Speechno summary yeteess-as2011.07755Tencent2 citesNov 16, 2020
- NovWPD++: An Improved Neural Beamformer for Simultaneous Speech Separation and Dereverberationno summary yeteess-as2011.09162Tencent0 citesNov 18, 2020
- NovUsing Synthetic Audio to Improve The Recognition of Out-Of-Vocabulary Words in End-To-End ASR Systemsno summary yeteess-as2011.11564Amazon1 citesNov 23, 2020
- NovSynth2Aug: Cross-domain speaker recognition with TTS synthesized speechno summary yeteess-as2011.11818DeepMind1 citesNov 24, 2020
- OctTraining Strategies to Handle Missing Modalities for Audio-Visual Expression Recognitionno summary yeteess-as2010.00734Amazon1 citesOct 2, 2020
- OctTowards Data-efficient Modeling for Wake Word Spottingno summary yeteess-as2010.06659Amazon1 citesOct 13, 2020
- OctPushing the Limits of Semi-Supervised Learning for Automatic Speech Recognitionno summary yeteess-as2010.10504Google Research200 citesOct 20, 2020
- OctReal-time Speech Frequency Bandwidth Extensionno summary yeteess-as2010.10677Google Research1 citesOct 21, 2020
- OctFastEmit: Low-latency Streaming ASR with Sequence-level Emission Regularizationno summary yeteess-as2010.11148Google Research6 citesOct 21, 2020
- OctConfidence Estimation for Attention-based Sequence-to-sequence Models for Speech Recognitionno summary yeteess-as2010.11428Google Research3 citesOct 22, 2020
- OctMicrosoft Speaker Diarization System for the VoxCeleb Speaker Recognition Challenge 2020no summary yeteess-as2010.11458Microsoft Research6 citesOct 22, 2020
- OctCascaded encoders for unifying streaming and non-streaming ASRno summary yeteess-as2010.14606Google Research5 citesOct 27, 2020
- OctReplay and Synthetic Speech Detection with Res2net Architectureno summary yeteess-as2010.15006Tencent18 citesOct 28, 2020
- OctOptimizing Short-Time Fourier Transform Parameters via Gradient Descentno summary yeteess-as2010.15049Adobe17 citesOct 28, 2020
- OctDirectional ASR: A New Paradigm for E2E Multi-Speaker Speech Recognition with Source Localizationno summary yeteess-as2011.00091Tencent4 citesOct 30, 2020
- SepWaveGrad: Estimating Gradients for Waveform Generationno summary yeteess-as2009.00713Google Research44 citesSep 2, 2020
- SepSAGRNN: Self-Attentive Gated RNN for Binaural Speaker Separation with Interaural Cue Preservationno summary yeteess-as2009.01381Meta / FAIR21 citesSep 2, 2020
- SepSEANet: A Multi-modal Speech Enhancement Networkno summary yeteess-as2009.02095Google Research6 citesSep 4, 2020
- SepVoiceFilter-Lite: Streaming Targeted Voice Separation for On-Device Speech Recognitionno summary yeteess-as2009.04323Google Research10 citesSep 9, 2020
- SepICASSP 2021 Deep Noise Suppression Challengeno summary yeteess-as2009.06122Microsoft Research23 citesSep 14, 2020
- SepDiffWave: A Versatile Diffusion Model for Audio Synthesisno summary yeteess-as2009.09761Baidu121 citesSep 21, 2020
- SepA consolidated view of loss functions for supervised deep learning-based speech enhancementno summary yeteess-as2009.12286Microsoft Research13 citesSep 25, 2020
- AugMultimodal Semi-supervised Learning Framework for Punctuation Prediction in Conversational Speechno summary yeteess-as2008.00702Amazon1 citesAug 3, 2020
- AugMusiCoder: A Universal Music-Acoustic Encoder Based on Transformersno summary yeteess-as2008.00781Alibaba14 citesAug 3, 2020
- AugA Spectral Energy Distance for Parallel Speech Synthesisno summary yeteess-as2008.01160Google Research20 citesAug 3, 2020
- AugLearning to Denoise Historical Musicno summary yeteess-as2008.02027Google Research2 citesAug 5, 2020
- AugData balancing for boosting performance of low-frequency classes in Spoken Language Understandingno summary yeteess-as2008.02603Amazon4 citesAug 6, 2020
- AugA Machine of Few Words -- Interactive Speaker Recognition with Reinforcement Learningno summary yeteess-as2008.03127DeepMind0 citesAug 7, 2020
- AugControllable Neural Prosody Synthesisno summary yeteess-as2008.03388Adobe2 citesAug 7, 2020
- AugSpeech Driven Talking Face Generation from a Single Image and an Emotion Conditionno summary yeteess-as2008.03592Microsoft Research9 citesAug 8, 2020
- AugVariable frame rate-based data augmentation to handle speaking-style variability for automatic speaker verificationno summary yeteess-as2008.03616Amazon1 citesAug 8, 2020
- AugDisentangled Multidimensional Metric Learning for Music Similarityno summary yeteess-as2008.03720Adobe0 citesAug 9, 2020
- AugSubword Regularization: An Analysis of Scalability and Generalization for End-to-End Automatic Speech Recognitionno summary yeteess-as2008.04034Amazon0 citesAug 10, 2020
- AugA Perceptually-Motivated Approach for Low-Complexity, Real-Time Enhancement of Fullband Speechno summary yeteess-as2008.04259Amazon3 citesAug 10, 2020
- AugPoCoNet: Better Speech Enhancement with Frequency-Positional Embeddings, Semi-Supervised Conversational Data, and Biased Lossno summary yeteess-as2008.04470Amazon9 citesAug 11, 2020
- AugInvestigation of End-To-End Speaker-Attributed ASR for Continuous Multi-Talker Recordingsno summary yeteess-as2008.04546Microsoft Research3 citesAug 11, 2020
- AugOnline Automatic Speech Recognition with Listen, Attend and Spell Modelno summary yeteess-as2008.05514Apple19 citesAug 12, 2020
- AugTextual Echo Cancellationno summary yeteess-as2008.06006Google Research7 citesAug 13, 2020
- AugData augmentation and loss normalization for deep noise suppressionno summary yeteess-as2008.06412Microsoft Research13 citesAug 14, 2020
- AugAdaptation Algorithms for Neural Network-Based Speech Recognition: An Overviewno summary yeteess-as2008.06580Microsoft Research73 citesAug 14, 2020
- AugADL-MVDR: All deep learning MVDR beamformer for target speech separationno summary yeteess-as2008.06994Tencent7 citesAug 16, 2020
- AugImproving Tail Performance of a Deliberation E2E ASR Model Using a Large Text Corpusno summary yeteess-as2008.10491Google Research7 citesAug 24, 2020
- AugLearned Transferable Architectures Can Surpass Hand-Designed Architectures for Large Scale Speech Recognitionno summary yeteess-as2008.11589Tencent1 citesAug 25, 2020
- AugParallel Rescoring with Transformer for Streaming On-Device Speech Recognitionno summary yeteess-as2008.13093Google Research1 citesAug 30, 2020
- JulResNeXt and Res2Net Structures for Speaker Verificationno summary yeteess-as2007.02480Microsoft Research13 citesJul 6, 2020
- JulTERA: Self-Supervised Learning of Transformer Encoder Representation for Speechno summary yeteess-as2007.06028Amazon313 citesJul 12, 2020
- JulSudo rm -rf: Efficient Networks for Universal Audio Source Separationno summary yeteess-as2007.06833Adobe108 citesJul 14, 2020
- JulStreaming ResLSTM with Causal Mean Aggregation for Device-Directed Utterance Detectionno summary yeteess-as2007.09245Amazon4 citesJul 17, 2020
- JulDNN No-Reference PSTN Speech Quality Predictionno summary yeteess-as2007.14598Microsoft Research1 citesJul 29, 2020
- JunUnsupervised Sound Separation Using Mixture Invariant Trainingno summary yeteess-as2006.12701Google Research91 citesJun 23, 2020
- MayContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Contextno summary yeteess-as2005.03191Google Research72 citesMay 7, 2020
- MayRNN-T Models Fail to Generalize to Out-of-Domain Audio: Causes and Solutionsno summary yeteess-as2005.03271Google Research12 citesMay 7, 2020
- MayStreaming keyword spotting on mobile devicesno summary yeteess-as2005.06720Google Research112 citesMay 14, 2020
- MayAudio ALBERT: A Lite BERT for Self-supervised Learning of Audio Representationno summary yeteess-as2005.08575Amazon39 citesMay 18, 2020
- MayTransferring Source Style in Non-Parallel Voice Conversionno summary yeteess-as2005.09178Tencent2 citesMay 19, 2020
- MayImproved Noisy Student Training for Automatic Speech Recognitionno summary yeteess-as2005.09629Google Research228 citesMay 19, 2020
- MayImproving Proper Noun Recognition in End-to-End ASR By Customization of the MWER Loss Criterionno summary yeteess-as2005.09756Google Research1 citesMay 19, 2020
- MayTraining Keyword Spotting Models on Non-IID Data with Federated Learningno summary yeteess-as2005.10406Google Research13 citesMay 21, 2020
- MayDynamic Sparsity Neural Networks for Automatic Speech Recognitionno summary yeteess-as2005.10627Google Research2 citesMay 16, 2020
- MayGlottal source estimation robustness: A comparison of sensitivity of voice source estimation techniquesno summary yeteess-as2005.11682Amazon1 citesMay 24, 2020
- MayModality Dropout for Improved Performance-driven Talking Facesno summary yeteess-as2005.13616Apple1 citesMay 27, 2020
- AprF0-consistent many-to-many non-parallel voice conversion via conditional autoencoderno summary yeteess-as2004.07370Adobe97 citesApr 15, 2020
- AprMatchboxNet: 1D Time-Channel Separable Convolutional Neural Network Architecture for Speech Commands Recognitionno summary yeteess-as2004.08531NVIDIA85 citesApr 18, 2020
- AprLanguage-agnostic Multilingual Modelingno summary yeteess-as2004.09571Google Research2 citesApr 20, 2020
- AprViSQOL v3: An Open Source Production Ready Objective Speech and Audio Metricno summary yeteess-as2004.09584Google Research2 citesApr 20, 2020
- AprJukebox: A Generative Model for Musicno summary yeteess-as2005.00341OpenAI109 citesApr 30, 2020
- MarEnhancing End-to-End Multi-channel Speech Separation via Spatial Feature Learningno summary yeteess-as2003.03927Tencent4 citesMar 9, 2020
- MarMulti-modal Multi-channel Target Speech Separationno summary yeteess-as2003.07032Tencent110 citesMar 16, 2020
- MarHigh-Accuracy and Low-Latency Speech Recognition with Two-Head Contextual Layer Trajectory LSTM Modelno summary yeteess-as2003.07482Microsoft Research2 citesMar 17, 2020
- MarHybrid Autoregressive Transducer (hat)no summary yeteess-as2003.07705Google Research3 citesMar 12, 2020
- MarDeliberation Model Based Two-Pass End-to-End Speech Recognitionno summary yeteess-as2003.07962Google Research1 citesMar 17, 2020
- MarSpeech Quality Factors for Traditional and Neural-Based Low Bit Rate Vocodersno summary yeteess-as2003.11882Google Research0 citesMar 26, 2020
- FebTensor-to-Vector Regression for Multi-channel Speech Enhancement based on Tensor-Train Networkno summary yeteess-as2002.00544Tencent3 citesFeb 3, 2020
- FebBOFFIN TTS: Few-Shot Speaker Adaptation by Bayesian Optimizationno summary yeteess-as2002.01953Amazon57 citesFeb 4, 2020
- FebTransformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Lossno summary yeteess-as2002.02562Google Research27 citesFeb 7, 2020
- FebFully-hierarchical fine-grained prosody modeling for interpretable speech synthesisno summary yeteess-as2002.03785Google Research7 citesFeb 6, 2020
- FebGenerating diverse and natural text-to-speech samples using a quantized fine-grained VAE and auto-regressive prosody priorno summary yeteess-as2002.03788Google Research11 citesFeb 6, 2020
- FebMultimodal active speaker detection and virtual cinematography for video conferencingno summary yeteess-as2002.03977Meta / FAIR0 citesFeb 10, 2020
- FebA Comparison of Pooling Methods on LSTM Models for Rare Acoustic Event Classificationno summary yeteess-as2002.06279Salesforce2 citesFeb 14, 2020
- FebImputer: Sequence Modelling via Imputation and Dynamic Programmingno summary yeteess-as2002.08926Google Research23 citesFeb 20, 2020
- FebWavesplit: End-to-End Speech Separation by Speaker Clusteringno summary yeteess-as2002.08933Google Research73 citesFeb 20, 2020
- FebMulti-label Sound Event Retrieval Using a Deep Learning-based Siamese Structure with a Pairwise Presence Matrixno summary yeteess-as2002.09026Microsoft Research0 citesFeb 20, 2020
- FebEfficient Trainable Front-Ends for Neural Speech Enhancementno summary yeteess-as2002.09286Amazon0 citesFeb 20, 2020
- FebA Density Ratio Approach to Language Model Fusion in End-To-End Automatic Speech Recognitionno summary yeteess-as2002.11268Google Research2 citesFeb 26, 2020
- FebTowards Learning a Universal Non-Semantic Representation of Speechno summary yeteess-as2002.12764Google Research124 citesFeb 25, 2020
- JanAudio-visual Recognition of Overlapped speech for the LRS2 datasetno summary yeteess-as2001.01656Tencent10 citesJan 6, 2020
- JanA Differentiable Perceptual Audio Metric Learned from Just Noticeable Differencesno summary yeteess-as2001.04460Adobe5 citesJan 13, 2020
- JanA Memory Augmented Architecture for Continuous Speaker Identification in Meetingsno summary yeteess-as2001.05118Microsoft Research3 citesJan 15, 2020
- JanLow-rank Gradient Approximation For Memory-Efficient On-device Training of Deep Neural Networkno summary yeteess-as2001.08885Google Research0 citesJan 24, 2020
- JanSemi-supervised ASR by End-to-end Self-trainingno summary yeteess-as2001.09128Salesforce7 citesJan 24, 2020
- JanData Techniques For Online End-to-end Speech Recognitionno summary yeteess-as2001.09221Salesforce4 citesJan 24, 2020
- JanWeighted Speech Distortion Losses for Neural-network-based Real-time Speech Enhancementno summary yeteess-as2001.10601Microsoft Research17 citesJan 28, 2020
- JanUnsupervised Pre-training of Bidirectional Speech Encoders via Masked Reconstructionno summary yeteess-as2001.10603Salesforce1 citesJan 28, 2020
- JanEnvironment-aware Reconfigurable Noise Suppressionno summary yeteess-as2001.10718Meta / FAIR0 citesJan 29, 2020
- JanLattice-based Improvements for Voice Triggering Using Graph Neural Networksno summary yeteess-as2001.10822Apple1 citesJan 25, 2020
- JanTraining Keyword Spotters with Limited and Synthesized Speech Datano summary yeteess-as2002.01322Google Research7 citesJan 31, 2020
- JanDetecting Emotion Primitives from Speech and their use in discerning Categorical Emotionsno summary yeteess-as2002.01323Apple1 citesJan 31, 2020
2019
41- DecDeep Contextualized Acoustic Representations For Semi-Supervised Speech Recognitionno summary yeteess-as1912.01679Amazon150 citesDec 3, 2019
- DecAudio-attention discriminative language model for ASR rescoringno summary yeteess-as1912.03363Amazon0 citesDec 6, 2019
- DecSpecAugment on Large Scale Datasetsno summary yeteess-as1912.05533Google Research4 citesDec 11, 2019
- DecPersonalization of End-to-end Speech Recognition On Mobile Devices For Named Entitiesno summary yeteess-as1912.09251Google Research3 citesDec 14, 2019
- NovPredicting word error rate for reverberant speechno summary yeteess-as1911.00566Microsoft Research0 citesNov 1, 2019
- NovASVspoof 2019: A large-scale public database of synthesized, converted and replayed speechno summary yeteess-as1911.01601Google Research441 citesNov 5, 2019
- NovSpatial Attention for Far-field Speech Recognition with Deep Beamforming Neural Networksno summary yeteess-as1911.02115Meta / FAIR0 citesNov 5, 2019
- NovRecurrent Neural Network Transducer for Audio-Visual Speech Recognitionno summary yeteess-as1911.04890DeepMind15 citesNov 8, 2019
- OctDual-path RNN: efficient long sequence modeling for time-domain single-channel speech separationno summary yeteess-as1910.06379Microsoft Research50 citesOct 14, 2019
- OctQuartzNet: Deep Automatic Speech Recognition with 1D Time-Channel Separable Convolutionsno summary yeteess-as1910.10261NVIDIA31 citesOct 22, 2019
- OctRecognizing long-form speech using streaming end-to-end modelsno summary yeteess-as1910.11455Google Research4 citesOct 24, 2019
- OctLearning audio representations via phase predictionno summary yeteess-as1910.11910Google Research10 citesOct 25, 2019
- OctMixup-breakdown: a consistency training method for improving generalization of speech separation modelsno summary yeteess-as1910.13253Tencent2 citesOct 28, 2019
- OctDFSMN-SAN with Persistent Memory Model for Automatic Speech Recognitionno summary yeteess-as1910.13282Tencent5 citesOct 28, 2019
- OctEnd-to-end Microphone Permutation and Number Invariant Multi-channel Speech Separationno summary yeteess-as1910.14104Microsoft Research7 citesOct 30, 2019
- OctFast acoustic scattering using convolutional neural networksno summary yeteess-as1911.01802Microsoft Research7 citesOct 30, 2019
- SepEvaluating Long-form Text-to-Speech: Comparing the Ratings of Sentences and Paragraphsno summary yeteess-as1909.03965Google Research1 citesSep 9, 2019
- SepGenerative Speech Enhancement Based on Cloned Networksno summary yeteess-as1909.04776Google Research1 citesSep 10, 2019
- SepLarge-Scale Multilingual Speech Recognition with a Streaming End-to-End Modelno summary yeteess-as1909.05330Google Research4 citesSep 11, 2019
- SepSams-Net: A Sliced Attention-based Neural Network for Music Source Separationno summary yeteess-as1909.05746Tencent2 citesSep 12, 2019
- SepAn Investigation Into On-device Personalization of End-to-end Automatic Speech Recognition Modelsno summary yeteess-as1909.06678Google Research22 citesSep 14, 2019
- SepDiPCo -- Dinner Party Corpusno summary yeteess-as1909.13447Amazon7 citesSep 30, 2019
- AugExploiting semi-supervised training through a dropout regularization in end-to-end speech recognitionno summary yeteess-as1908.05227Adobe0 citesAug 8, 2019
- AugSalient Speech Representations Based on Cloned Networksno summary yeteess-as1908.07045Google Research1 citesAug 19, 2019
- AugConnecting and Comparing Language Model Interpolation Techniquesno summary yeteess-as1908.09738Google Research0 citesAug 26, 2019
- JulImproving Performance of End-to-End ASR on Numeric Sequencesno summary yeteess-as1907.01372Google Research5 citesJul 1, 2019
- JulSpeech bandwidth extension with WaveNetno summary yeteess-as1907.04927DeepMind4 citesJul 5, 2019
- JulTeach an all-rounder with experts in different domainsno summary yeteess-as1907.05698Tencent0 citesJul 9, 2019
- JunThe Second DIHARD Diarization Challenge: Dataset, task, and baselinesno summary yeteess-as1906.07839Baidu27 citesJun 18, 2019
- JunLipper: Synthesizing Thy Speech using Multi-View Lipreadingno summary yeteess-as1907.01367Adobe7 citesJun 28, 2019
- MayImproving Opus Low Bit Rate Quality with Neural Speech Synthesisno summary yeteess-as1905.04628Google Research4 citesMay 12, 2019
- MaySpeaker-Independent Speech-Driven Visual Speech Synthesis using Domain-Adapted Acoustic Modelsno summary yeteess-as1905.06860Apple1 citesMay 15, 2019
- AprParrotron: An End-to-End Speech-to-Speech Conversion Model and its Applications to Hearing-Impaired Speech and Speech Separationno summary yeteess-as1904.04169Google Research19 citesApr 8, 2019
- AprLow-Latency Speaker-Independent Continuous Speech Separationno summary yeteess-as1904.06478Microsoft Research2 citesApr 13, 2019
- MarSinging voice conversion with non-parallel datano summary yeteess-as1903.04124Snap0 citesMar 11, 2019
- MarNon-intrusive speech quality assessment using neural networksno summary yeteess-as1903.06908Microsoft Research4 citesMar 16, 2019
- MarA Real-Time Wideband Neural Vocoder at 1.6 kb/s Using LPCNetno summary yeteess-as1903.12087Amazon6 citesMar 28, 2019
- FebA spelling correction model for end-to-end speech recognitionno summary yeteess-as1902.07178Google Research2 citesFeb 19, 2019
- FebUtterance-level end-to-end language identification using attention-based CNN-BLSTMno summary yeteess-as1902.07374Tencent3 citesFeb 20, 2019
- JanImproving noise robustness of automatic speech recognition via parallel data and teacher-student learningno summary yeteess-as1901.02348Amazon22 citesJan 5, 2019
- JanSelf-Attention Networks for Connectionist Temporal Classification in Speech Recognitionno summary yeteess-as1901.10055Amazon133 citesJan 22, 2019
2018
18- DecUnsupervised Speech Recognition via Segmental Empirical Output Distribution Matchingno summary yeteess-as1812.09323Tencent2 citesDec 23, 2018
- DecEnd-to-End Classification of Reverberant Rooms using DNNsno summary yeteess-as1812.09324Amazon4 citesDec 21, 2018
- NovTowards achieving robust universal neural vocodingno summary yeteess-as1811.06292Amazon9 citesNov 15, 2018
- OctRecognizing Overlapped Speech in Meetings: A Multichannel Separation Approach Using Neural Networksno summary yeteess-as1810.03655Microsoft Research101 citesOct 8, 2018
- OctLPCNet: Improving Neural Speech Synthesis Through Linear Predictionno summary yeteess-as1810.11846Google Research25 citesOct 28, 2018
- OctContextual Speech Recognition with Difficult Negative Training Examplesno summary yeteess-as1810.12170Google Research2 citesOct 29, 2018
- OctLow-Dimensional Bottleneck Features for On-Device Continuous Speech Recognitionno summary yeteess-as1811.00006Google Research0 citesOct 31, 2018
- SepCycle-Consistent Speech Enhancementno summary yeteess-as1809.02253Microsoft Research57 citesSep 6, 2018
- SepFrom Audio to Semantics: Approaches to end-to-end spoken language understandingno summary yeteess-as1809.09190Google Research14 citesSep 24, 2018
- AugDeep context: end-to-end contextual speech recognitionno summary yeteess-as1808.02480Google Research9 citesAug 7, 2018
- JulLearning Noise-Invariant Representations for Robust Speech Recognitionno summary yeteess-as1807.06610Amazon4 citesJul 17, 2018
- JulA Comparison of Techniques for Language Model Integration in Encoder-Decoder Speech Recognitionno summary yeteess-as1807.10857Google Research9 citesJul 27, 2018
- JunMulti-View Networks for Denoising of Arbitrary Numbers of Channelsno summary yeteess-as1806.05296Adobe0 citesJun 13, 2018
- AprA Novel Learnable Dictionary Encoding Layer for End-to-End Language Identificationno summary yeteess-as1804.00385Tencent1 citesApr 2, 2018
- AprAdversarial Teacher-Student Learning for Unsupervised Domain Adaptationno summary yeteess-as1804.00644Microsoft Research87 citesApr 2, 2018
- AprSpeaker-Invariant Training via Adversarial Learningno summary yeteess-as1804.00732Microsoft Research116 citesApr 2, 2018
- MarLinear networks based speaker adaptation for speech synthesisno summary yeteess-as1803.02445Alibaba4 citesMar 5, 2018
- MarDirectional emphasis in ambisonicsno summary yeteess-as1803.06718Google Research0 citesMar 18, 2018
2017
6- DecWavenet based low rate speech codingno summary yeteess-as1712.01120DeepMind2 citesDec 1, 2017
- DecMulti-Dialect Speech Recognition With A Single Sequence-To-Sequence Modelno summary yeteess-as1712.01541Google Research10 citesDec 5, 2017
- DecAn analysis of incorporating an external language model into a sequence-to-sequence modelno summary yeteess-as1712.01996Google Research11 citesDec 6, 2017
- NovDeep Long Short-Term Memory Adaptive Beamforming Networks For Multichannel Robust Speech Recognitionno summary yeteess-as1711.08016Microsoft Research103 citesNov 21, 2017
- OctGeneralized End-to-End Loss for Speaker Verificationno summary yeteess-as1710.10467Google Research28 citesOct 28, 2017
- OctSpeaker Diarization with LSTMno summary yeteess-as1710.10468Google Research2 citesOct 28, 2017