AI TOOLS / MARCH – APRIL 2026
Discord AI Bridge
A Discord bot that answers questions with a model running on my own machine.
- MY ROLE
- Solo project
- KEY RESULT
- Local inference moved off the event loop
- TEAM
- Built on my own.
How a request flows
- Discord!.ask command
- discord.pyEvent loop stays free
- Worker threadasyncio.to_thread()
- LM StudioQwen 3.5 9B, local
Why
I wanted to see what it takes to put a locally hosted language model behind a real chat interface, with no cloud API in between.
What I built
- 01
Connected discord.py to LM Studio’s OpenAI-compatible endpoint and a local Qwen 3.5 9B model, triggered by the !.ask command.
- 02
Moved the blocking HTTP call into a worker thread with asyncio.to_thread(), so the bot keeps handling other messages while a reply is generated.
- 03
Turned connection failures, HTTP errors, and unexpected responses into plain messages in the channel instead of silent failures.
The code
try:
# Run the blocking request in a thread so the bot stays responsive
response = await asyncio.to_thread(requests.post, API_URL, json=payload)
response.raise_for_status()
ai_response = response.json()["choices"][0]["message"]["content"]
await message.channel.send(
ai_response, allowed_mentions=discord.AllowedMentions(everyone=False)
)
except requests.exceptions.ConnectionError:
await message.channel.send("⚠️ Error: LM Studio server is not running.")
except requests.exceptions.HTTPError as e:
await message.channel.send(f"⚠️ Server Error: {e.response.status_code}")Outcome
People in the server can ask the local model a question with one command, and the bot stays responsive while it answers.
The trade-off
I kept the requests library and wrapped it in asyncio.to_thread() instead of switching to an async HTTP client like aiohttp. It fixed the blocking with one line and code I already understood. The cost is a thread per request and no way to cancel one once it starts.
What I learned
Async code only stays async if every slow call respects it. One blocking request can freeze the whole bot.
What I'd fix next
Add a timeout to the model request (today a hung model server holds a thread forever), and split replies longer than Discord’s 2,000-character message limit.