Centific logo
Centific
Posted 75 days agoVerified live 7h ago

AI Research Engineer- Speech 1

Brief overview

Remote
MastersOr in progress
$150k–$160k/yrStated range
2+ yrsMinimum
86 H-1B approvalsDept. of Labor
7 green cardsCertified filings
PythonPyTorchGPU-accelerated trainingSpeech and audio signal processingAcoustic modelingAudio representationsLarge Language ModelsTransformersState Space Models (SSMs)Instruction tuningAlignment methodsAdapter-based integrationCross-modal attentionAudio-text fusionLarge Audio Language ModelsSpeech-to-Speech systemsNeural audio codecs

Visa sponsorship history

4 years sponsoring, last filed FY2026

Data powered by U.S. Department of Labor. This does not guarantee sponsorship for this specific role.
86H-1B approved
98%approval rate
27new H-1B hires
7PERM certified
$110,750median wage / yr
H-1B Petition ApprovalsVisas USCIS actually granted: the strongest sign the company sponsors.
202318
202437
202521
202610
LCA Certified ApplicationsAn early filing step, not a visa approval: it signals intent, not confirmed sponsorship.
20238
20245
20253
20265
Green Card (PERM) FilingsCertified green card filings: a long-term commitment to international hires.
20233
20241
20252
20261
Top sponsored roles
Software Development EngineerData ScientistSoftware Development Engineer in TestData AnalystSystems Analyst II
Sponsored employees from
IndiaChina

Job description

Summary

Centific is a frontier AI data foundry that empowers clients with innovative AI solutions. They are seeking an AI Research Engineer specializing in Speech/Audio to develop next-generation audio AI technologies, focusing on Large Audio Language Models and Speech-to-Speech systems.

Responsibilities

  • Design, develop, and deploy Large Audio Language Models (LALMs) capable of native audio understanding, reasoning, and generation
  • Build Large Audio Reasoning Models that perform complex chain-of-thought reasoning over speech and audio inputs, including medical, technical, and conversational domains
  • Contribute to Speech-to-Speech (S2S) system development, including speech understanding, dialogue management, and speech synthesis components
  • Research and implement alignment mechanisms between speech encoders and LLM backbones using lightweight adapters, LoRA, and efficient fine-tuning strategies
  • Design efficient speech tokenization and temporal compression techniques suitable for long-form audio reasoning and multi-turn spoken dialogue
  • Build comprehensive evaluation frameworks for audio reasoning capabilities, including benchmarks for speech QA, audio understanding, and reasoning accuracy
  • Optimize inference pipelines for low-latency, streaming applications in speech systems
  • Collaborate with cross-functional teams to transfer research innovations into production systems and customer-facing applications
  • Contribute to technical documentation, research write-ups, and publications at top-tier venues (NeurIPS, ICML, ACL, Interspeech)

Skills

  • Master's degree (required) or Ph.D. (preferred) in Computer Science, Electrical Engineering, or a related field with a focus on speech, audio ML, or multimodal learning
  • 2+ years of industry or applied research experience in speech/audio AI, Large Language Models, or multimodal systems
  • Demonstrated applied research contributions through publications, patents, or shipped products in speech/audio AI or LLMs
  • Strong proficiency in Python and PyTorch, with hands-on experience in GPU-accelerated training for large-scale models
  • Solid understanding of speech and audio signal processing, acoustic modeling, and audio representations
  • Working knowledge of modern LLM architectures (Transformers, SSMs) and training paradigms including instruction tuning and alignment methods
  • Familiarity with modality alignment techniques: adapter-based integration, cross-modal attention, or audio-text fusion methods
  • Strong experimentation habits: clean code, systematic ablations, reproducibility, and clear technical communication
  • Publication record at top-tier venues (NeurIPS, ICML, ICLR, ACL, Interspeech, ICASSP) in audio language models, speech reasoning, or multimodal learning
  • Hands-on experience building or fine-tuning Large Audio Language Models (e.g., Qwen-Audio, SALMONN, LTU, Gemini Audio)
  • Experience with speech representation pretraining (HuBERT, Wav2Vec 2.0, Whisper, WavLM) and discrete speech tokenization
  • Familiarity with Speech-to-Speech components: neural audio codecs (EnCodec, SoundStream), vocoders, or speech synthesis systems
  • Experience with audio reasoning benchmarks (AIR-Bench, MMAU, AudioBench) or building evaluation harnesses for audio QA
  • Hands-on experience with distributed training (FSDP, DeepSpeed) and inference optimization (ONNX, TensorRT, quantization)
  • Familiarity with speech frameworks such as ESPnet, SpeechBrain, NVIDIA NeMo, or Fairseq
  • Experience with multilingual speech systems, code-switching, or domain adaptation for specialized applications (medical, legal, technical)
  • Background in evaluating safety, bias, hallucination, or adversarial robustness in audio language models

Qualifications

Must Haves

  • Master's degree (required) or Ph.D. (preferred) in Computer Science, Electrical Engineering, or a related field with a focus on speech, audio ML, or multimodal learning
  • 2+ years of industry or applied research experience in speech/audio AI, Large Language Models, or multimodal systems
  • Demonstrated applied research contributions through publications, patents, or shipped products in speech/audio AI or LLMs
  • Strong proficiency in Python and PyTorch, with hands-on experience in GPU-accelerated training for large-scale models
  • Solid understanding of speech and audio signal processing, acoustic modeling, and audio representations
  • Working knowledge of modern LLM architectures (Transformers, SSMs) and training paradigms including instruction tuning and alignment methods
  • Familiarity with modality alignment techniques: adapter-based integration, cross-modal attention, or audio-text fusion methods
  • Strong experimentation habits: clean code, systematic ablations, reproducibility, and clear technical communication

Nice to Haves

  • Publication record at top-tier venues (NeurIPS, ICML, ICLR, ACL, Interspeech, ICASSP) in audio language models, speech reasoning, or multimodal learning
  • Hands-on experience building or fine-tuning Large Audio Language Models (e.g., Qwen-Audio, SALMONN, LTU, Gemini Audio)
  • Experience with speech representation pretraining (HuBERT, Wav2Vec 2.0, Whisper, WavLM) and discrete speech tokenization
  • Familiarity with Speech-to-Speech components: neural audio codecs (EnCodec, SoundStream), vocoders, or speech synthesis systems
  • Experience with audio reasoning benchmarks (AIR-Bench, MMAU, AudioBench) or building evaluation harnesses for audio QA
  • Hands-on experience with distributed training (FSDP, DeepSpeed) and inference optimization (ONNX, TensorRT, quantization)
  • Familiarity with speech frameworks such as ESPnet, SpeechBrain, NVIDIA NeMo, or Fairseq
  • Experience with multilingual speech systems, code-switching, or domain adaptation for specialized applications (medical, legal, technical)
  • Background in evaluating safety, bias, hallucination, or adversarial robustness in audio language models

Benefits

  • Competitive compensation package with comprehensive benefits
  • Opportunity to work on cutting-edge Large Audio Language Models and audio reasoning research with real-world impact
  • Collaboration with experienced applied scientists and engineers in speech and multimodal AI
  • Support for publications at top-tier conferences and professional development
  • Access to state-of-the-art GPU infrastructure for training large-scale audio models
  • Flexible work arrangements with hybrid/remote options

More jobs like this