Research Scientist
Overview
I. Job Details
Experience – 4+years in AI/ML, with strong hands-on experience in speech/Audio/Voice AI
Who we’re looking for
We’re looking for Research Scientists with strong hands-on experience in Voice AI. Depending on your strengths, this might mean independently building, evaluating, optimizing, and deploying production-grade speech/audio models – or formulating research problems, designing rigorous experiments, developing novel approaches, and driving projects toward state-of-the-art results and publications/patents where appropriate. Most Research Scientists here do a mix of both.
Candidates may apply against any one (or more) of the following problem tracks, based on their interest and prior experience.
Across every track, you will
- Drive AI research and innovation in Speech/Audio AI
- Translate research questions into measurable hypotheses; design rigorous experiments and ablations; analyze model behavior and failure modes; reproduce and build on relevant papers
- Take research from prototype to production-grade, deployable systems
- Define appropriate objective and subjective evaluation metrics, and run benchmarking, error analysis, and statistical comparisons to know when a model is actually better – not just different • Collaborate cross-functionally with engineering, product, and DSP/embedded teams
- Bring strong problem-solving ability – debug, iterate, and optimize models under real-world constraints
- Stay current with SOTA literature and bring new techniques into the team; publish or patent where appropriate
Track 1: Automatic Speech Recognition (ASR)
- Work on end-to-end ASR architectures (Transformer, Conformer, RNN-T) for streaming speech recognition
- Work with speech foundation models (Whisper-style architectures) and self-supervised models (HuBERT, wav2vec 2.0, WavLM) for representation learning and fine-tuning
- Own the end-to-end ASR pipeline – preprocessing, feature extraction, acoustic modeling, and decoding/language modeling
- Research streaming and causal architectures under strict latency constraints
- Improve robustness under noisy, multi-accent, and low-resource conditions – starting with Indian-English / regional Indian accents, with a roadmap to scale to international accents and languages
- Research accent-invariant / accent-normalized representations – removing accent information from learned embeddings to improve generalization across accents and geographies
- Study representation quality and robustness across speakers, accents, recording conditions, and domains
- Evaluated on: WER, CER, RTF, latency, memory
- Preferred background: prior work on CTC/RNN-T, self-supervised pretraining, or Whisper finetuning
Track 2: Text-to-Speech (TTS)
- Work on flow-matching / diffusion-based TTS systems (in the spirit of CosyVoice, ZipVoice, VITS, and similar SOTA architectures) for real-time, low-compute synthesis
- Experience with duration-controlled, time-aligned TTS using forced alignment (e.g., MFA) and Monotonic Alignment Search (MAS) for multi-speaker settings
- Research controllable accent adaptation using speaker, accent, linguistic, prosodic, and/or style representations – including decoder-side conditioning and representation-level approaches – starting with Indian-English and US-English accent pairs and scaling to international accents • Research controllable duration, timing, rhythm, F0, and prosody modeling for natural accent rendering
- Improve naturalness, prosody, and voice-cloning quality
- Optimize inference for real-time, streaming synthesis on constrained hardware
- Evaluated on: MOS, speaker similarity, intelligibility, prosody, RTF
- Preferred background: prior work with neural vocoders, diffusion/flow-matching TTS, or zeroshot voice cloning
Track 3: Speech Enhancement / Noise Suppression
- Develop time-domain, spectral-domain, and hybrid speech enhancement models for real-time applications
- Work across both single-microphone and multi-microphone / multi-channel scenarios, including beamforming-informed approaches
- Improve performance in low-SNR, reverberant, and non-stationary noise conditions
- Evaluate using objective and perceptual metrics under challenging real-world acoustic conditions
- Optimize for low latency and low compute for edge / embedded deployment
- Evaluated on: SI-SDR, PESQ, STOI, DNSMOS, latency, compute
- Preferred background: prior work on single- or multi-channel speech enhancement, beamforming, or DNS Challenge-style benchmarks
Track 4: Model Optimization • Work on quantization (including quantization-aware training), pruning, and knowledge distillation for speech/audio models
- Optimize inference pipelines (e.g., ONNX / TensorRT-class toolchains) for low-latency, lowpower edge and embedded deployment
- Profile end-to-end pipelines across CPU/GPU/NPU targets to identify model- and system-level bottlenecks – the neural network isn’t always the actual latency bottleneck
- Benchmark and profile models across hardware platforms, balancing latency, throughput, and memory trade-offs
• Preferred background: prior work with model compression, quantization-aware training, or on device inference
Also of interest
We’re also excited about adjacent problems – target speaker extraction, accent conversion, source separation, bandwidth extension, and packet loss concealment, Speech Translation. If you’ve worked on any of these, tell us about it.
Core Skills / Qualifications Required:
- Strong programming skills in Python and AI/ML frameworks (PyTorch, TensorFlow)
- Strong ML/DL fundamentals, including CNNs, RNNs, LSTMs, and Transformers applied to speech/audio
- Good understanding of signal processing, linear algebra, optimization techniques, statistics, and pattern recognition
- Minimum 4 years of hands-on experience in the Audio/Speech/Voice AI domain
- Ability to independently design and run experiments; strong debugging and analytical skills
- Self-motivated to learn, explore new areas, and work independently
- Hands-on experience with Generative AI concepts, including GANs, Flow Matching, and Diffusion Models.
Nice to have:
- Working knowledge of C/C++ and CUDA (required for Track 4: Model Optimization)
- Experience with ONNX/TensorRT, Kaldi, ESPnet, NeMo, or Hugging Face
- Experience with embedded/edge deployment
- Publications, patents, or open-source contributions in speech/audio
We value strong fundamentals, research thinking, experimental rigor, and problem-solving ability over familiarity with any specific framework or model architecture.
Qualification: Bachelor’s / Master’s / PhD in AI/ML, CSE, ECE, or related field
II. About Us
Company Profile
Meeami Technologies is a leader in Audio AI, Noise Cancellation, Speaker ID and other Audio Analytics technologies. Meeami has over 75+ patents granted and pending in this Audio AI & analytics technologies. Meeami, a pioneer with over 20+ years of experience in Audio solutions, is a spin-off of the former media processing and real-time communications group of Imagination Technologies, is the recognized leader in IP Communications and Voice IoT technology platforms for voice, video and messaging applications. To see how Meeami is helping top-tier OEM, IC, Call Centers and carrier customers, with embedded software, mobile apps and end-to-end communications solutions
Website
https://www.meeamitech.com
Location
Meeami Technologies Private Limited, plot no #37, 3rd floor, Hitech City road, Madhapur, Hyderabad , Telangana , 500081
Contact Email
recruitment@meeamitech.com & hr@meeamitech.com
III. Recruitment Procedure
Round 1
Technical Discussion
Round 2
Technical Assignment & Coding – problem formulation, implementation, experimentation, analysis, and presentation
Round 3
Deep Technical Discussion
Round 4
HR & Managerial Round