Multimodal AI Engineer-FT

An AI company is looking for a Multimodal AI Engineer. This is an IN-PERSON position in Manhattan, New York.

Responsibilities:

  • Build Multimodal AI Systems
    • Design and develop AI systems that process text, speech, audio, and visual inputs simultaneously
    • Build real-time multimodal inference pipelines for live interview and meeting assistance
    • Develop cross-modal reasoning systems combining audio, visual, and textual understanding
  • Speech & Audio AI Engineering
    • Build low-latency speech-to-text pipelines using technologies like Whisper, Deepgram, or AssemblyAI
    • Implement speaker diarization, sentiment analysis, emotion detection, and speech understanding systems
    • Optimize streaming audio processing and live transcription performance
  • Visual & Document AI
    • Develop systems for OCR, document parsing, screen analysis, and visual context extraction
    • Build AI pipelines for understanding resumes, presentations, coding screenshots, and interview materials
    • Integrate visual understanding into AI reasoning workflows
  • Multimodal Model Integration
    • Integrate advanced multimodal foundation models such as GPT-4o, Gemini, Claude Vision, and open-source multimodal systems
    • Build multimodal RAG pipelines combining text, image, speech, and document retrieval
    • Improve contextual understanding across multiple AI modalities
  • Real-Time AI Optimization
    • Optimize multimodal inference for low-latency production performance
    • Build concurrent processing systems using streaming and asynchronous architectures
    • Improve CPU, GPU, and memory utilization across AI workloads

Requirements:

  • Strong proficiency in Python
  • Experience with PyTorch, TensorFlow, torchaudio, librosa, or speech processing frameworks
  • Expertise with speech-to-text systems including Whisper, Deepgram, Google Speech, or AssemblyAI
  • Knowledge of computer vision, OCR, and document AI systems
  • Experience integrating multimodal models such as GPT-4o, Gemini, or LLaVA
  • Familiarity with WebSockets, streaming systems, and real-time AI pipelines
  • Experience with LangChain, LlamaIndex, and vector databases for multimodal RAG systems

Salary: $160,000 – $230,000

Hours: Full-Time

Zip Code: 10003

 Job ID:  #69736MAE

Interested in this job? Become a client to learn more.

Scroll to Top