AI TOOLS / MARCH – APRIL 2026

Discord AI Bridge

A Discord bot that answers questions with a model running on my own machine.

MY ROLE
Solo project
KEY RESULT
Local inference moved off the event loop
TEAM
Built on my own.

How a request flows

  1. Discord!.ask command
  2. discord.pyEvent loop stays free
  3. Worker threadasyncio.to_thread()
  4. LM StudioQwen 3.5 9B, local

Why

I wanted to see what it takes to put a locally hosted language model behind a real chat interface, with no cloud API in between.

What I built

  1. 01

    Connected discord.py to LM Studio’s OpenAI-compatible endpoint and a local Qwen 3.5 9B model, triggered by the !.ask command.

  2. 02

    Moved the blocking HTTP call into a worker thread with asyncio.to_thread(), so the bot keeps handling other messages while a reply is generated.

  3. 03

    Turned connection failures, HTTP errors, and unexpected responses into plain messages in the channel instead of silent failures.

The code

discord_ai_bridge/main.pyView on GitHub
try:
    # Run the blocking request in a thread so the bot stays responsive
    response = await asyncio.to_thread(requests.post, API_URL, json=payload)
    response.raise_for_status()
    ai_response = response.json()["choices"][0]["message"]["content"]

    await message.channel.send(
        ai_response, allowed_mentions=discord.AllowedMentions(everyone=False)
    )
except requests.exceptions.ConnectionError:
    await message.channel.send("⚠️ Error: LM Studio server is not running.")
except requests.exceptions.HTTPError as e:
    await message.channel.send(f"⚠️ Server Error: {e.response.status_code}")
The blocking call runs in a worker thread, so the event loop keeps serving other messages.

Outcome

People in the server can ask the local model a question with one command, and the bot stays responsive while it answers.

The trade-off

I kept the requests library and wrapped it in asyncio.to_thread() instead of switching to an async HTTP client like aiohttp. It fixed the blocking with one line and code I already understood. The cost is a thread per request and no way to cancel one once it starts.

What I learned

Async code only stays async if every slow call respects it. One blocking request can freeze the whole bot.

What I'd fix next

Add a timeout to the model request (today a hung model server holds a thread forever), and split replies longer than Discord’s 2,000-character message limit.

Go deeper

How one blocking call froze my Discord bot4 min read

Built with

  • Python
  • discord.py
  • LM Studio
  • REST API
  • asyncio