Portfolio - Fishchenko

02. Babbily

About the Project

Multi-functional aggregator of AI models, combining text, audio, video, and multimodal neural networks in a single user-friendly interface.

Mission

Rethink the architecture for AI interaction via APIs. Instead of outdated logic, we implemented a plugin-based approach supporting streaming and faster switching between models. We simplified thread storage structures and expanded functionality.

Challenges

  1. High complexity in synchronizing numerous models with different output formats
  2. Building unified logic for interface and backend communication
  3. Rapid functionality scaling without compromising stability

Integrations and plugins

Connected over 25 models, including:
  1. Text: GPT-4.0, GPT-4.5, Claude 3.5, Claude 3.7, Gemini 1.5 Pro, Gemini 2.5 Pro, LLaMA 3, Command R+, Qwen Max, and others
  2. Audio: ElevenLabs (TTS), Mochi 2.0 (TTS or ASR)
  3. Images: Google Imagen 3
  4. Multimodal: Gemini, OpenAI, Sonar Pro Reasoning, DeepSeek V3, and more

Development stages

  1. Legacy Product Analysis: identified current architecture limitations
  2. Full Migration: transitioned from API Connector to a plugin system as per client request
  3. Plugin Management System: developed a unified management interface
  4. Model Integrations: connected over 15 models within 1 month, including TTS, SST, video and audio generators, image generators, and document analyzers
  5. Streaming Tests: tested real-time output streaming and system stability
  6. Thread Logic Update: improved saving and restoring threads for scalability

Key AI Resources

  1. OpenAI: GPT 3.5, GPT 4.0, GPT 4.5, O3 Mini
  2. Anthropic (Claude): Claude 3.5 Haiku, Claude 3.7 Sonnet
  3. Google (Gemini): Gemini 1.5 Pro, Gemini 2.0 Flash, Gemini 2.5 Pro, Imagen 3
  4. Sonar: Sonar Pro, Sonar Pro Reasoning, Sonar Deep Research
  5. DeepSeek: DeepSeek R1, DeepSeek V3
  6. Cohere: Command R Plus, Command R+
  7. Alibaba (Qwen): Qwen Max, Qwen QWQ 32
  8. Mistral: Mistral Pxtral Large
  9. xAI (Grok): Grok
  10. Reka / Together AI: Together, Flux 1.1 Pro
  11. Deepgram: speech recognition
  12. ElevenLabs: speech generation (TTS/STT)
  13. Runway: video generation (Gen-3 Alpha)
  14. Custom: Mochi 2.0, Boba 2.0

Link

AI SaaS
Made on
Tilda