134 papers
cs sd
0/02026
5- JulMulTTiPop: A Multitrack Transcription Dataset for Pop Musicno summary yetcs-sd2607.08756CMU0 citesJul 9, 2026
- JunAugCodec: A Low-Bitrate Disentangled Neural Speech Codec via Data Augmentationno summary yetcs-sd2606.21893Amazon0 citesJun 20, 2026
- JunInstructFX2FX: A Multi-Turn Text-to-Effect System for Sequential Audio Effect Refinementno summary yetcs-sd2606.22005Berkeley0 citesJun 20, 2026
- JunAdaptive Perturbation Selection for Contrastive Audio Decodingno summary yetcs-sd2607.00247Google Research0 citesJun 30, 2026
- FebFine-grained Soundscape Control for Augmented Hearingno summary yetcs-sd2603.00395UW0 citesFeb 28, 2026
2024
6- OctExpGest: Expressive Speaker Generation Using Diffusion Model and Hybrid Audio-Text Guidanceno summary yetcs-sd2410.09396Tencent4 citesOct 12, 2024
- AugThe evolution of inharmonicity and noisiness in contemporary popular musicno summary yetcs-sd2408.08127Sony0 citesAug 15, 2024
- AugImproving Generalization of Speech Separation in Real-World Scenarios: Strategies in Simulation, Optimization, and Evaluationno summary yetcs-sd2408.16126Adobe0 citesAug 28, 2024
- MayLook Once to Hear: Target Speech Hearing with Noisy Examplesno summary yetcs-sd2405.06289Microsoft Research23 citesMay 10, 2024
- FebObjective and subjective evaluation of speech enhancement methods in the UDASE task of the 7th CHiME challengeno summary yetcs-sd2402.01413Google Research14 citesFeb 2, 2024
- FebAdvancing Audio Fingerprinting Accuracy Addressing Background Noise and Distortion Challengesno summary yetcs-sd2402.13957Google Research29 citesFeb 21, 2024
2023
6- DecCreating New Voices using Normalizing Flowsno summary yetcs-sd2312.14569Amazon10 citesDec 22, 2023
- NovSemantic Hearing: Programming Acoustic Scenes with Binaural Hearablesno summary yetcs-sd2311.00320Microsoft Research28 citesNov 1, 2023
- NovDecoupling and Interacting Multi-Task Learning Network for Joint Speech and Accent Recognitionno summary yetcs-sd2311.07062Tencent15 citesNov 13, 2023
- SepText-Only Domain Adaptation for End-to-End Speech Recognition through Down-Sampling Acoustic Representationno summary yetcs-sd2309.02459Tencent0 citesSep 4, 2023
- AugConformer-based Target-Speaker Automatic Speech Recognition for Single-Channel Audiono summary yetcs-sd2308.05218NVIDIA16 citesAug 9, 2023
- MarWESPER: Zero-shot and Realtime Whisper to Normal Voice Conversion for Whisper-based Speech Interactionsno summary yetcs-sd2303.01639Sony22 citesMar 3, 2023
2022
5- SepImproving Choral Music Separation through Expressive Synthesized Data from Sampled Instrumentsno summary yetcs-sd2209.02871Tencent3 citesSep 7, 2022
- AugChronological Self-Training for Real-Time Speaker Diarizationno summary yetcs-sd2208.03393Google Research0 citesAug 5, 2022
- JulReLyMe: Improving Lyric-to-Melody Generation by Incorporating Lyric-Melody Relationshipsno summary yetcs-sd2207.05688Alibaba12 citesJul 12, 2022
- AprLeveraging Phone Mask Training for Phonetic-Reduction-Robust E2E Uyghur Speech Recognitionno summary yetcs-sd2204.00819Tencent9 citesApr 2, 2022
- FebSummary On The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challengeno summary yetcs-sd2202.03647Alibaba24 citesFeb 8, 2022
2021
26- OctAuto-DSP: Learning to Optimize Acoustic Echo Cancellersno summary yetcs-sd2110.04284Adobe2 citesOct 8, 2021
- SepStructure-Enhanced Pop Music Generation via Harmony-Aware Learningno summary yetcs-sd2109.06441Tencent22 citesSep 14, 2021
- SepRendering Spatial Sound for Interoperable Experiences in the Audio Metaverseno summary yetcs-sd2109.12471Meta / FAIR4 citesSep 26, 2021
- SepEmergency Vehicles Audio Detection and Localization in Autonomous Drivingno summary yetcs-sd2109.14797Baidu12 citesSep 30, 2021
- JulDance2Music: Automatic Dance-driven Music Generationno summary yetcs-sd2107.06252Adobe8 citesJul 13, 2021
- JunImproving On-Screen Sound Separation for Open-Domain Videos with Audio-Visual Self-Attentionno summary yetcs-sd2106.09669Google Research3 citesJun 17, 2021
- MayEnd-to-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddingsno summary yetcs-sd2105.02096Google Research0 citesMay 5, 2021
- MaySeparate but Together: Unsupervised Federated Learning for Speech Enhancement from Non-IID Datano summary yetcs-sd2105.04727Adobe14 citesMay 11, 2021
- MayThe Benefit Of Temporally-Strong Labels In Audio Event Classificationno summary yetcs-sd2105.07031Google Research7 citesMay 14, 2021
- MayMondegreen: A Post-Processing Solution to Speech Recognition Error Correction for Voice Search Queriesno summary yetcs-sd2105.09930Google Research2 citesMay 20, 2021
- MayDIVE: End-to-end Speech Diarization via Iterative Speaker Embeddingno summary yetcs-sd2105.13802Google Research0 citesMay 28, 2021
- AprAISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenariono summary yetcs-sd2104.03603Microsoft Research5 citesApr 8, 2021
- AprSpectrogram Inpainting for Interactive Generation of Instrument Soundsno summary yetcs-sd2104.07519Sony5 citesApr 15, 2021
- AprAdaSpeech 2: Adaptive Text to Speech with Untranscribed Datano summary yetcs-sd2104.09715Microsoft Research0 citesApr 20, 2021
- AprMultimodal Self-Supervised Learning of General Audio Representationsno summary yetcs-sd2104.12807Google Research20 citesApr 26, 2021
- MarSymbolic Music Generation with Diffusion Modelsno summary yetcs-sd2103.16091Google Research9 citesMar 30, 2021
- FebMulti-Task Self-Supervised Pre-Training for Music Classificationno summary yetcs-sd2102.03229Amazon4 citesFeb 5, 2021
- FebLightSpeech: Lightweight and Fast Text to Speech with Neural Architecture Searchno summary yetcs-sd2102.04040Microsoft Research4 citesFeb 8, 2021
- FebA Multi-View Approach To Audio-Visual Speaker Verificationno summary yetcs-sd2102.06291Meta / FAIR5 citesFeb 11, 2021
- FebContrastive Unsupervised Learning for Speech Emotion Recognitionno summary yetcs-sd2102.06357Amazon6 citesFeb 12, 2021
- FebMulti-Channel Speech Enhancement using Graph Neural Networksno summary yetcs-sd2102.06934Meta / FAIR7 citesFeb 13, 2021
- FebMBNet: MOS Prediction for Synthesized Speech with Mean-Bias Networkno summary yetcs-sd2103.00110Microsoft Research5 citesFeb 27, 2021
- JanGeneralized Spatio-Temporal RNN Beamformer for Target Speech Separationno summary yetcs-sd2101.01280Tencent2 citesJan 4, 2021
- JanHypothesis Stitcher for End-to-End Speaker-attributed ASR on Long-form Multi-talker Recordingsno summary yetcs-sd2101.01853Microsoft Research1 citesJan 6, 2021
- JanInterspeech 2021 Deep Noise Suppression Challengeno summary yetcs-sd2101.01902Microsoft Research15 citesJan 6, 2021
- JanLEAF: A Learnable Frontend for Audio Classificationno summary yetcs-sd2101.08596Google Research30 citesJan 21, 2021
2020
31- NovSound Event Detection and Separation: a Benchmark on Desed Synthetic Soundscapesno summary yetcs-sd2011.00801Adobe3 citesNov 2, 2020
- NovWhat's All the FUSS About Free Universal Sound Separation Data?no summary yetcs-sd2011.00803Adobe9 citesNov 2, 2020
- NovInto the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen Soundsno summary yetcs-sd2011.01143Google Research9 citesNov 2, 2020
- NovA Two-Stage Approach to Device-Robust Acoustic Scene Classificationno summary yetcs-sd2011.01447Tencent45 citesNov 3, 2020
- NovDESNet: A Multi-channel Network for Simultaneous Speech Dereverberation, Enhancement and Separationno summary yetcs-sd2011.02131Microsoft Research1 citesNov 4, 2020
- NovBW-EDA-EEND: Streaming End-to-End Neural Speaker Diarization for a Variable Number of Speakersno summary yetcs-sd2011.02678Amazon2 citesNov 5, 2020
- NovDual Application of Speech Enhancement for Automatic Speech Recognitionno summary yetcs-sd2011.03840Meta / FAIR2 citesNov 7, 2020
- NovFRILL: A Non-Semantic Speech Embedding for Mobile Devicesno summary yetcs-sd2011.04609Google Research0 citesNov 9, 2020
- NovCommunication-Cost Aware Microphone Selection For Neural Speech Enhancement with Ad-hoc Microphone Arraysno summary yetcs-sd2011.07348Adobe1 citesNov 14, 2020
- NovAdversarial Training for Multi-domain Speaker Recognitionno summary yetcs-sd2011.08623Tencent1 citesNov 17, 2020
- NovStreaming end-to-end multi-talker speech recognitionno summary yetcs-sd2011.13148Microsoft Research41 citesNov 26, 2020
- OctTransformer Transducer: One Model Unifying Streaming and Non-streaming Speech Recognitionno summary yetcs-sd2010.03192Google Research30 citesOct 7, 2020
- OctNon-Attentive Tacotron: Robust and Controllable Neural TTS Synthesis Including Unsupervised Duration Modelingno summary yetcs-sd2010.04301Google Research73 citesOct 8, 2020
- OctAI Song Contest: Human-AI Co-Creation in Songwritingno summary yetcs-sd2010.05388Google Research15 citesOct 12, 2020
- OctMicAugment: One-shot Microphone Style Transferno summary yetcs-sd2010.09658Google Research1 citesOct 19, 2020
- OctPrediction of Object Geometry from Acoustic Scattering Using Convolutional Neural Networksno summary yetcs-sd2010.10691Microsoft Research2 citesOct 21, 2020
- OctParallel Tacotron: Non-Autoregressive and Controllable TTSno summary yetcs-sd2010.11439Google Research21 citesOct 22, 2020
- OctImproving Streaming Automatic Speech Recognition With Non-Streaming Model Distillation On Unsupervised Datano summary yetcs-sd2010.12096Google Research1 citesOct 22, 2020
- OctLearning Fine-Grained Cross Modality Excitement for Speech Emotion Recognitionno summary yetcs-sd2010.12733Baidu1 citesOct 24, 2020
- OctNon-Autoregressive Transformer ASR with CTC-Enhanced Decoder Inputno summary yetcs-sd2010.15025Tencent4 citesOct 28, 2020
- OctDNSMOS: A Non-Intrusive Perceptual Objective Speech Quality metric to evaluate Noise Suppressorsno summary yetcs-sd2010.15258Microsoft Research58 citesOct 28, 2020
- OctLearning Audio Embeddings with User Listening Data for Content-based Music Recommendationno summary yetcs-sd2010.15389Tencent4 citesOct 29, 2020
- OctRespireNet: A Deep Neural Network for Accurately Detecting Abnormal Lung Sounds in Limited Data Settingno summary yetcs-sd2011.00196Microsoft Research0 citesOct 31, 2020
- JulImproving Sound Event Detection In Domestic Environments Using Sound Separationno summary yetcs-sd2007.03932Adobe28 citesJul 8, 2020
- JunEnd-to-End Adversarial Text-to-Speechno summary yetcs-sd2006.03575Google Research33 citesJun 5, 2020
- MayAddressing Missing Labels in Large-Scale Sound Event Recognition Using a Teacher-Student Framework With Loss Maskingno summary yetcs-sd2005.00878Google Research23 citesMay 2, 2020
- AprImproving Perceptual Quality of Drum Transcription with the Expanded Groove MIDI Datasetno summary yetcs-sd2004.00188Google Research14 citesApr 1, 2020
- FebFully Learnable Front-End for Multi-Channel Acoustic Modeling using Semi-Supervised Learningno summary yetcs-sd2002.00125Amazon1 citesFeb 1, 2020
- FebRobust Multi-channel Speech Recognition using Frequency Aligned Networkno summary yetcs-sd2002.02520Amazon0 citesFeb 6, 2020
- JanCURE Dataset: Ladder Networks for Audio Event Classificationno summary yetcs-sd2001.03896Microsoft Research0 citesJan 12, 2020
- JanContinuous speech separation: dataset and analysisno summary yetcs-sd2001.11482Microsoft Research11 citesJan 30, 2020
2019
30- DecWaveFlow: A Compact Flow-based Model for Raw Audiono summary yetcs-sd1912.01219Baidu30 citesDec 3, 2019
- DecPitchNet: Unsupervised Singing Voice Conversion with Pitch Adversarial Networkno summary yetcs-sd1912.01852Tencent1 citesDec 4, 2019
- DecVoice Conversion for Whispered Speech Synthesisno summary yetcs-sd1912.05289Amazon29 citesDec 11, 2019
- DecEncoding Musical Style with Transformer Autoencodersno summary yetcs-sd1912.05537Google Research9 citesDec 10, 2019
- DecDeep Audio Priorno summary yetcs-sd1912.10292Adobe6 citesDec 21, 2019
- NovCoincidence, Categorization, and Consolidation: Learning to Recognize Sounds with Minimal Supervisionno summary yetcs-sd1911.05894Google Research0 citesNov 14, 2019
- NovScene-Aware Audio Rendering via Deep Acoustic Analysisno summary yetcs-sd1911.06245Adobe38 citesNov 14, 2019
- NovImproving Universal Sound Separation Using Sound Classificationno summary yetcs-sd1911.07951Google Research71 citesNov 18, 2019
- NovSequential Multi-Frame Neural Beamforming for Speech Separation and Enhancementno summary yetcs-sd1911.07953Microsoft Research5 citesNov 18, 2019
- NovMusic Source Separation in the Waveform Domainno summary yetcs-sd1911.13254Meta / FAIR185 citesNov 27, 2019
- OctEnd-to-End Multi-Task Denoising for the Joint Optimization of Perceptual Speech Metricsno summary yetcs-sd1910.10707Google Research2 citesOct 23, 2019
- OctSecost: Sequential co-supervision for large scale weakly labeled audio event detectionno summary yetcs-sd1910.11789Meta / FAIR8 citesOct 25, 2019
- OctMellotron: Multispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokensno summary yetcs-sd1910.11997NVIDIA17 citesOct 26, 2019
- SepDemucs: Deep Extractor for Music Sources with extra unlabeled data remixedno summary yetcs-sd1909.01174Meta / FAIR58 citesSep 3, 2019
- SepImpulse Response Data Augmentation and Deep Neural Networks for Blind Room Acoustic Parameter Estimationno summary yetcs-sd1909.03642Adobe1 citesSep 9, 2019
- SepA scalable noisy speech dataset and online subjective test frameworkno summary yetcs-sd1909.08050Microsoft Research16 citesSep 17, 2019
- SepHigh Fidelity Speech Synthesis with Adversarial Networksno summary yetcs-sd1909.11646DeepMind104 citesSep 25, 2019
- JulImproving Reverberant Speech Training Using Diffuse Acoustic Simulationno summary yetcs-sd1907.03988Tencent34 citesJul 9, 2019
- JulThe Bach Doodle: Approachable music composition with machine learning at scaleno summary yetcs-sd1907.06637Google Research42 citesJul 14, 2019
- JunAudio tagging with noisy labels and minimal supervisionno summary yetcs-sd1906.02975Google Research22 citesJun 7, 2019
- MayUniversal Sound Separationno summary yetcs-sd1905.03330Google Research8 citesMay 8, 2019
- MayLearning to Groove with Inverse Sequence Transformationsno summary yetcs-sd1905.06118Google Research1 citesMay 14, 2019
- AprTowards Generalized Speech Enhancement with Generative Adversarial Networksno summary yetcs-sd1904.03418Amazon6 citesApr 6, 2019
- AprOn Acoustic Modeling for Broadband Beamformingno summary yetcs-sd1904.08971Amazon0 citesApr 18, 2019
- AprRealizing Petabyte Scale Acoustic Modelingno summary yetcs-sd1904.10584Amazon5 citesApr 24, 2019
- AprAdversarial Speaker Verificationno summary yetcs-sd1904.12406Microsoft Research62 citesApr 29, 2019
- AprPerforming Structured Improvisations with pre-trained Deep Learning Modelsno summary yetcs-sd1904.13285Google Research1 citesApr 30, 2019
- AprDeep Learning for Audio Signal Processingno summary yetcs-sd1905.00078Google Research818 citesApr 30, 2019
- FebGANSynth: Adversarial Neural Audio Synthesisno summary yetcs-sd1902.08710Google Research70 citesFeb 23, 2019
- JanLearning Sound Event Classifiers from Web Audio with Noisy Labelsno summary yetcs-sd1901.01189Google Research13 citesJan 4, 2019
2018
14- DecInverSynth: Deep Estimation of Synthesizer Parameter Configurations from Audio Signalsno summary yetcs-sd1812.06349Microsoft Research1 citesDec 15, 2018
- NovMulti-View Networks For Multi-Channel Audio Classificationno summary yetcs-sd1811.01251Adobe1 citesNov 3, 2018
- NovSDR - half-baked or well done?no summary yetcs-sd1811.02508Microsoft Research16 citesNov 6, 2018
- NovExploring Tradeoffs in Models for Low-latency Speech Enhancementno summary yetcs-sd1811.07030Google Research6 citesNov 16, 2018
- NovDifferentiable Consistency Constraints for Improved Deep Speech Enhancementno summary yetcs-sd1811.08521Google Research6 citesNov 20, 2018
- OctPhasebook and Friends: Leveraging Discrete Representations for Source Separationno summary yetcs-sd1810.01395Google Research79 citesOct 2, 2018
- OctSING: Symbol-to-Instrument Neural Generatorno summary yetcs-sd1810.09785Meta / FAIR26 citesOct 23, 2018
- OctEnabling Factorized Piano Music Modeling and Generation with the MAESTRO Datasetno summary yetcs-sd1810.12247Google Research149 citesOct 29, 2018
- OctWaveGlow: A Flow-based Generative Network for Speech Synthesisno summary yetcs-sd1811.00002NVIDIA75 citesOct 31, 2018
- SepSelf-Supervised Generation of Spatial Audio for 360 Videono summary yetcs-sd1809.02587Adobe79 citesSep 7, 2018
- AugThis Time with Feeling: Learning Expressive Musical Performanceno summary yetcs-sd1808.03715DeepMind23 citesAug 10, 2018
- MayConvolutional-Recurrent Neural Networks for Speech Enhancementno summary yetcs-sd1805.00579Microsoft Research13 citesMay 2, 2018
- AprLooking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separationno summary yetcs-sd1804.03619Google Research572 citesApr 10, 2018
- JanWaveform Modeling and Generation Using Hierarchical Recurrent Neural Networks for Speech Bandwidth Extensionno summary yetcs-sd1801.07910Baidu65 citesJan 24, 2018
2017
7- DecOn Using Backpropagation for Speech Texture Generation and Voice Conversionno summary yetcs-sd1712.08363Google Research2 citesDec 22, 2017
- NovKnowledge Transfer from Weakly Labeled Audio using Convolutional Neural Network for Sound Events and Scenesno summary yetcs-sd1711.01369Meta / FAIR13 citesNov 4, 2017
- NovUnsupervised Learning of Semantic Audio Representationsno summary yetcs-sd1711.02209Google Research5 citesNov 6, 2017
- NovExploring Speech Enhancement with Generative Adversarial Networks for Robust Speech Recognitionno summary yetcs-sd1711.05747Google Research21 citesNov 15, 2017
- OctOnsets and Frames: Dual-Objective Piano Transcriptionno summary yetcs-sd1710.11153Google Research43 citesOct 30, 2017
- SepDeep Learning Techniques for Music Generation -- A Surveyno summary yetcs-sd1709.01620Sony199 citesSep 5, 2017
- MayEnd-to-end Source Separation with Adaptive Front-Endsno summary yetcs-sd1705.02514Adobe15 citesMay 6, 2017
2016
3- SepCNN Architectures for Large-Scale Audio Classificationno summary yetcs-sd1609.09430Google Research14 citesSep 29, 2016
- JunFast, Compact, and High Quality LSTM-RNN Based Statistical Parametric Speech Synthesizers for Mobile Devicesno summary yetcs-sd1606.06061Google Research12 citesJun 20, 2016
- MaySpeech Enhancement In Multiple-Noise Conditions using Deep Neural Networksno summary yetcs-sd1605.02427Microsoft Research19 citesMay 9, 2016