An AI company is looking for a Multimodal AI Engineer. This is an IN-PERSON position in Manhattan, New York.
Responsibilities:
- Build Multimodal AI Systems
- Design and develop AI systems that process text, speech, audio, and visual inputs simultaneously
- Build real-time multimodal inference pipelines for live interview and meeting assistance
- Develop cross-modal reasoning systems combining audio, visual, and textual understanding
- Speech & Audio AI Engineering
- Build low-latency speech-to-text pipelines using technologies like Whisper, Deepgram, or AssemblyAI
- Implement speaker diarization, sentiment analysis, emotion detection, and speech understanding systems
- Optimize streaming audio processing and live transcription performance
- Visual & Document AI
- Develop systems for OCR, document parsing, screen analysis, and visual context extraction
- Build AI pipelines for understanding resumes, presentations, coding screenshots, and interview materials
- Integrate visual understanding into AI reasoning workflows
- Multimodal Model Integration
- Integrate advanced multimodal foundation models such as GPT-4o, Gemini, Claude Vision, and open-source multimodal systems
- Build multimodal RAG pipelines combining text, image, speech, and document retrieval
- Improve contextual understanding across multiple AI modalities
- Real-Time AI Optimization
- Optimize multimodal inference for low-latency production performance
- Build concurrent processing systems using streaming and asynchronous architectures
- Improve CPU, GPU, and memory utilization across AI workloads
Requirements:
- Strong proficiency in Python
- Experience with PyTorch, TensorFlow, torchaudio, librosa, or speech processing frameworks
- Expertise with speech-to-text systems including Whisper, Deepgram, Google Speech, or AssemblyAI
- Knowledge of computer vision, OCR, and document AI systems
- Experience integrating multimodal models such as GPT-4o, Gemini, or LLaVA
- Familiarity with WebSockets, streaming systems, and real-time AI pipelines
- Experience with LangChain, LlamaIndex, and vector databases for multimodal RAG systems
Salary: $160,000 – $230,000
Hours: Full-Time
Zip Code: 10003
Job ID: #69736MAE