Multi-functional aggregator of AI models, combining text, audio, video, and multimodal neural networks in a single user-friendly interface.
Mission
Rethink the architecture for AI interaction via APIs. Instead of outdated logic, we implemented a plugin-based approach supporting streaming and faster switching between models. We simplified thread storage structures and expanded functionality.
Challenges
High complexity in synchronizing numerous models with different output formats
Building unified logic for interface and backend communication
Rapid functionality scaling without compromising stability
Integrations and plugins
Connected over 25 models, including:
Text: GPT-4.0, GPT-4.5, Claude 3.5, Claude 3.7, Gemini 1.5 Pro, Gemini 2.5 Pro, LLaMA 3, Command R+, Qwen Max, and others
Audio: ElevenLabs (TTS), Mochi 2.0 (TTS or ASR)
Images: Google Imagen 3
Multimodal: Gemini, OpenAI, Sonar Pro Reasoning, DeepSeek V3, and more
Development stages
Legacy Product Analysis: identified current architecture limitations
Full Migration: transitioned from API Connector to a plugin system as per client request
Plugin Management System: developed a unified management interface
Model Integrations: connected over 15 models within 1 month, including TTS, SST, video and audio generators, image generators, and document analyzers
Streaming Tests: tested real-time output streaming and system stability
Thread Logic Update: improved saving and restoring threads for scalability
Key AI Resources
OpenAI: GPT 3.5, GPT 4.0, GPT 4.5, O3 Mini
Anthropic (Claude): Claude 3.5 Haiku, Claude 3.7 Sonnet