About the Project
Multi-functional aggregator of AI models, combining text, audio, video, and multimodal neural networks in a single user-friendly interface.
Mission
Rethink the architecture for AI interaction via APIs. Instead of outdated logic, we implemented a plugin-based approach supporting streaming and faster switching between models. We simplified thread storage structures and expanded functionality.
Challenges
- High complexity in synchronizing numerous models with different output formats
- Building unified logic for interface and backend communication
- Rapid functionality scaling without compromising stability
Integrations and plugins
Connected over 25 models, including:
- Text: GPT-4.0, GPT-4.5, Claude 3.5, Claude 3.7, Gemini 1.5 Pro, Gemini 2.5 Pro, LLaMA 3, Command R+, Qwen Max, and others
- Audio: ElevenLabs (TTS), Mochi 2.0 (TTS or ASR)
- Images: Google Imagen 3
- Multimodal: Gemini, OpenAI, Sonar Pro Reasoning, DeepSeek V3, and more
Development stages
- Legacy Product Analysis: identified current architecture limitations
- Full Migration: transitioned from API Connector to a plugin system as per client request
- Plugin Management System: developed a unified management interface
- Model Integrations: connected over 15 models within 1 month, including TTS, SST, video and audio generators, image generators, and document analyzers
- Streaming Tests: tested real-time output streaming and system stability
- Thread Logic Update: improved saving and restoring threads for scalability
Key AI Resources
- OpenAI: GPT 3.5, GPT 4.0, GPT 4.5, O3 Mini
- Anthropic (Claude): Claude 3.5 Haiku, Claude 3.7 Sonnet
- Google (Gemini): Gemini 1.5 Pro, Gemini 2.0 Flash, Gemini 2.5 Pro, Imagen 3
- Sonar: Sonar Pro, Sonar Pro Reasoning, Sonar Deep Research
- DeepSeek: DeepSeek R1, DeepSeek V3
- Cohere: Command R Plus, Command R+
- Alibaba (Qwen): Qwen Max, Qwen QWQ 32
- Mistral: Mistral Pxtral Large
- xAI (Grok): Grok
- Reka / Together AI: Together, Flux 1.1 Pro
- Deepgram: speech recognition
- ElevenLabs: speech generation (TTS/STT)
- Runway: video generation (Gen-3 Alpha)
- Custom: Mochi 2.0, Boba 2.0
