Category
Audio agents
1,474 Audio AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
1,474 agents · ranked by popularity · refine in the directory →
Top 100 Audio agents
A generative speech model for daily dialogue.
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
💬 Open source machine learning framework to automate text- and voice-based conversations: NLU, dialogue management, connect to Slack, Facebook, and more - Create chatbots and voice assistants
🧠 Leon is your open-source personal assistant.
Open-source framework for conversational voice AI agents
Very low latency speech to text, intent recognition, and text to speech, for building voice agents and interfaces
AI Agent Engineering Platform built on an Open Source TypeScript AI Agent Framework
Open Vision Agents by Stream. Build voice and vision agents quickly with any model or video provider. Uses Stream's edge network for ultra-low latency.
🤖 wukong-robot 是一个简单、灵活、优雅的中文语音对话机器人/智能音箱项目,支持ChatGPT多轮对话能力,还可能是首个支持脑机交互的开源智能音箱项目。
⚙️ All-in-One menu bar app, hide 💻MacBook Pro's notch, dark mode, AirPods, Shortcuts
A multi-function Discord bot
Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.
Voice-to-text dictation app with local (Nvidia Parakeet/Whisper) and cloud models (BYOK). Privacy-first and available cross-platform.
Mac app for crushing tech interviews with AI
The most awesome list about bots ⭐️🤖
SimpleMem: Efficient Lifelong Memory for LLM Agents — Text & Multimodal
Instantly generate AI-powered subtitles on your device. Works standalone or connects to DaVinci Resolve.
faster_whisper GUI with PySide6
World's first AI meeting copilot → The Invisible Companion for Work + Life
💬 SpeechGPT is a web application that enables you to converse with ChatGPT.
🤖️ Cross-platform AI language practice app (跨平台AI语言练习应用)
Voice AI SDK is a reusable Android library that gives any app a full voice-driven AI conversation pipeline in minutes. Voice Assistant + Android Voide AI + SDK + MVVM + Kotlin
Real-time AI assistant for Meta Ray-Ban smart glasses -- voice + vision + agentic actions via Gemini Live and OpenClaw
The Open Source Alternative to Cluely - A lightning-fast, privacy-first AI assistant that works seamlessly during meetings, interviews, and conversations without anyone knowing. Built with Tauri for native performance, just 10MB. Completely undetectable in video calls, screen shares, and recordings.
🔊 Awesome list for Whisper — an open-source AI-powered speech recognition system developed by OpenAI
Rasa Core is now part of the Rasa repo: An open source machine learning framework to automate text-and voice-based conversations
ElevenLabs UI is a component library and custom registry built on top of shadcn/ui to help you build multimodal agents faster.
LLM Agent Framework in ComfyUI includes MCP sever, Omost,GPT-sovits, ChatTTS,GOT-OCR2.0, and FLUX prompt nodes,access to Feishu,discord,and adapts to all llms with similar openai / aisuite interfaces, such as o1,ollama, gemini, grok, qwen, GLM, deepseek, kimi,doubao. Adapted to local llms, vlm, gguf such as llama-3.3 Janus-Pro, Linkage graphRAG
Baresip is a modular SIP User-Agent with audio and video support
EPUB to audiobook converter, optimized for Audiobookshelf, WebUI included
Text-To-Speech, RAG, and LLMs. All local!
Realtime Voice AI with 100+ Models on Arduino ESP32 with Secure Websockets and Edge Functions for AI Toys, Companions, and Devices
A local-first AI chat workspace for models, agents, skills, plugins, search, RAG, voice, memory, and artifacts.
:speech_balloon: Easy way to create conversation chats
百聆 是一个类似GPT-4o的语音对话机器人,通过ASR+LLM+TTS实现,集成DeepSeek R1等优秀大模型,接入openClaw,真正的个人语音助手,时延低至800ms,Mac等低配置也可运行,支持打断
"VideoAgent: All-in-One Agentic Framework for Video Understanding, Editing, and Remaking"
Talk to your Mac, query your docs, no cloud required. On-device voice AI + RAG
Using OpenAI's Whisper to automatically generate YouTube subtitles
the open-source virtual assistant for Ubuntu based Linux distributions
基于AI的工作效率提升工具(聊天、绘画、知识库、工作流、 MCP服务市场、语音输入输出、长期记忆) | Ai-based productivity tools (Chat,Draw,RAG,Workflow,MCP marketplace, ASR,TTS, Long-term memory etc)
An open-source AI Voice Agent that integrates with Asterisk/FreePBX using Audiosocket/RTP technology
A private logbook with a staff of personal AI assistants. Agents read what you record and propose what to do next — you approve the changes. End-to-end encrypted sync between your own devices — servers only ever see ciphertext. Local AI optional.
Interact with OpenAI's ChatGPT via Telegram and Voice.
AI Vtuber for Streaming on Youtube/Twitch
The open-source iOS app that's making quality voice transcription more accessible on mobile devices.
Build realtime AI voice agents using FastRTC for low-latency streaming, Superlinked for vector search, Twilio for live phone calls, and Runpod for scalable GPU deployment.
Open source voice dictation technology
🧸 Lobe Vidol - Making Virtual Idols Accessible for EveryOne
This app can now use Android, just like a human.
A complete voice AI frontend app for LiveKit Agents with Next.js
A talking LLM that runs on your own computer without needing the internet.
Real-time web cockpit for OpenClaw: voice conversations, agent automated kanban board, workspace/file control, sub-agent sessions, inline charts, and usage visibility.
The simplest and lowest-cost AI integration solution. If you like this project, please give it a Star~ | 最简单、最低成本的AI接入方案。喜欢本项目的话点个 Star 吧~
Sample Amazon Lex chat bot web interface
Convert any git repository into an engaging podcast
React Native binding of whisper.cpp.
Local voice chatbot for engaging conversations, powered by Ollama, Hugging Face Transformers, and Coqui TTS Toolkit
🎤 The easiest way to transcribe audio in Swift
Your CrewAI Powered Video Editing Assistant
Running speech to text model (whisper.cpp) in Unity3d on your local machine.
The most advanced, fully offline client-side AI suite on Android today.
Conversational voice AI agents
Local AI talk with a custom voice based on Zephyr 7B model. Uses RealtimeSTT with faster_whisper for transcription and RealtimeTTS with Coqui XTTS for synthesis.
Beginner-friendly AI conversation practice application
Example scripts for AI agents created with the Alan AI Platform.
Real-time transcription using faster-whisper
TypeScript AI agent framework: cognitive memory, runtime tool forging, multi-agent orchestration, 11 LLM providers.
An open source Ruby framework for text and voice chatbots. 🤖
Rust Agent Development Kit (ADK-Rust): Build AI agents in Rust with modular components for models, tools, memory, realtime voice, and more. ADK-Rust is a flexible framework for developing AI agents with simplicity and power. Model-agnostic, deployment-agnostic, optimized for frontier AI models. Includes support for real-time voice agents.
Open Source Voice Agent Platform
[CVPR 2025] Video Narration as Vocabulary & Video as Long Document
Make your meetings accessible to AI Agents
A conversational, AI device + software framework for companionship, entertainment, education, healthcare, IoT applications, and DIY robotics. Built with Python, NextJS, Arduino, ESP32, LLMs (GPT-4o), Deepgram STT and Azure TTS 🤖
One-stop handbook for building, deploying, and understanding LLM agents with 60+ skeletons, tutorials, ecosystem guides, and evaluation tools.
Your own personal voice assistant: Voice to Text to LLM to Speech, displayed in a web interface
Mac compatible Ollama Voice
Pocket-sized AI chatbot built using a RPI Zero 2w / 5
😆 A voice chatbot that can imitate your expression. OpenCV+Dlib+Live2D+Moments Recorder+Turing Robot+Iflytek IAT+Iflytek TTS
为 Bilibili、YouTube 及本地视频提供 AI 视频摘要和知识库.AI video summarizer and knowledge base for Bilibili, YouTube and local videos.
Production-grade Go SDK for building AI agents with long-term memory, knowledge retrieval, and voice — runnable as a library, a daemon, or a real-time pipeline.
OK | Every voice, every meme, every transaction makes $OK stronger and more vibrant. Powered by all of us—and now, AI agents. OK is not just OK — it’s $OK. $OK?
Collections of skills for building with ElevenLabs
smol-podcaster is your podcast production agent 🎙️
Open-source AI meeting copilot - real-time transcription, echo cancellation, and AI assistance. Captures system audio + mic, cancels echo via WebRTC AEC3, transcribes with Deepgram, and gives you Claude/OpenAI help during meetings. Runs locally on macOS and Windows.
Distill videos, PDFs, transcripts, and notes into source-backed teacher Agent Skills.
visionOS examples ⸺ Spatial Computing Accelerators for Apple Vision Pro
Stream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across various modality combinations.
A general framework for distilling human-created multimodal resources into reusable, executable skills that AI agents can browse, compose, and run, validated across diverse domains including web, PowerPoint, Excel, Blender, CAD, Unreal Engine 5, and REAPER-based music production.
How to use OpenAIs Whisper to transcribe and diarize audio files
⚡ Edgen: Local, private GenAI server alternative to OpenAI. No GPU required. Run AI models locally: LLMs (Llama2, Mistral, Mixtral...), Speech-to-text (whisper) and many others.
Dulus Ai — Free Agentic AI, Making Gemini web cappable of running bash commands in your terminal! [Gui, Web, Cli, Telegram, 2,000 MCP, 100K Skills . LiteLLM (100+ providers), local models via Ollama, /lang in 34 languages, Mesa Redonda, voice, OCR, MemPalace, embedded sandbox OS.] Buy $Dulus
say - command line tool for voice and video calling
Tiny truly local voice-activated LLM Agent that runs on a Raspberry Pi
Pipecat voice AI agents running locally on macOS
Jarvis is a voice-activated, conversational AI assistant powered by a local LLM (Qwen via Ollama). It listens for a wake word, processes spoken commands using a local language model with LangChain, and responds out loud via TTS. It supports tool-calling for dynamic functions like checking the current time.
AI-powered tool for real-time interview question transcription and response generation.
Any source (PDF, video, web, audio, text) to interactive learning package with quizzes, flashcards and spaced repetition. One command, 12-section study guide.
Install Tiledesk on your server using Helm for Kubernetes orchestration and Docker Compose for running multi-container Docker applications. Tiledesk provides an open-source solution comparable to Voiceflow, empowering you to create sophisticated LLM-enabled chatbots that seamlessly transition interactions to human agents when needed.