Back to all projects
AI Document Assistant
AI/ML

AI Document Assistant

Project Overview

An advanced full-stack web application that allows users to upload documents, audio, and video files and interact with an AI chatbot to ask questions and get summaries.

Project Details

Release DateJan 2025
TechnologyFastAPIReactFAISS

AI Document Assistant: Intelligent Multi-Modal Document Q&A

This project is an advanced full-stack web application designed to break down information silos by letting users upload PDF documents, MP3/WAV audio files, and MP4/MOV videos to interact with an AI-powered conversational assistant.

Full-Stack Architecture

  • **FastAPI & Python Backend:** Serves high-speed endpoints for file processing, text chunking, and AI inferences.
  • **Groq API & Llama 3.1 Orchestration:** Leverages Llama 3.1 model running on Groq hardware for ultra-low latency text summarization and contextual question-answering.
  • **Whisper & FFmpeg Audio Pipeline:** Decodes incoming audio/video files and extracts transcripts alongside word-level timestamps.
  • **Vector Retrieval:** Employs a FAISS vector database to execute semantic search and retrieve high-relevance chunks to inject into the LLM context.
  • **Modern React & MUI Frontend:** Responsive interface with drag-and-drop file uploads, media players mapping timestamps directly to playback segments, and real-time chat bubbles.
  • Key Engineering Challenges Resolved

  • **Large File Processing Bottlenecks:** Designed an asynchronous processing queue utilizing Redis as a broker and background tasks, providing progress bars and keeping the web server responsive.