# Voice Insight

Portfolio record by Eugene Livschitz

Canonical page: https://livytech.space/projects/voice-insight/

Audio-intelligence backend

Backend and AI-processing pipeline that turns raw customer-call recordings into validated, structured, analysis-ready inputs for downstream voice analysis.

## My contribution

Owned the Node.js orchestration layer, service contracts, request flow, validation, preprocessing, traceability, protected access, packaging, and multi-service deployment.

## The hard part

Making unreliable real-world audio and multiple specialized services behave like one stable customer API inside customer-controlled, sometimes offline or GPU-aware environments.

## What shipped

- Built upload validation, format and duration checks, silence detection, preprocessing, transcription handling, metadata enrichment, and cleanup paths.
- Implemented stereo/mono branching, channel handling, transcript cleanup and merging, readiness checks, and structured downstream payloads.
- Delivered protected Docker Compose deployments with client licensing, usage tracking, health endpoints, audit data, obfuscation, and environment-aware scripts.

## Technical details

- The orchestration layer integrates dedicated Python/FastAPI services for speaker profiling, transcription, and model inference without claiming ownership of those Python models.
- FFmpeg and ffprobe support media inspection and preprocessing before service handoff.
- Deployment work addressed offline dependencies, model packaging, GPU-aware configuration, diagnostics, and stable integration contracts.

## Technology

Node.js, Express, Docker, Docker Compose, FFmpeg, ffprobe, JWT, FastAPI integration, Transcription
