I'm building a service that I can connect to over HTTPS and interact with through a REPL-like conversational loop. The server should be able to control devices on my local network, including my Zigbee network. I'm wondering whether I should use a provider's Python SDK, a general HTTP client, or an existing agent framework. Calling an LLM through subprocesses feels fragile, so I'd like advice on a clean architecture for handling conversations, tool calls, and streaming responses. Security is also a concern because the service would have access to my LAN.
4 Answers
An agent framework such as PydanticAI can be useful here. Define the tools your assistant is allowed to call—such as reading Zigbee device state or turning a device on—and use typed schemas for their inputs and outputs. The model should request a tool call, your application should validate it, execute it, and then send the result back to the model. This is safer and more predictable than letting the model generate arbitrary shell commands.
Be very careful with the LAN access. An LLM-backed service with unrestricted network or shell permissions effectively becomes a remote administration interface. Put tools behind an allowlist, validate every argument, use least-privilege credentials, add authentication and audit logging, and consider isolating the service in a container or separate network segment. Don’t expose arbitrary command execution just to make the REPL feel like a terminal.
For a hosted model, avoid launching a command-line process from your server. Use the provider’s Python SDK or an HTTP client from an async endpoint, then stream the model’s output back to the client with server-sent events or WebSockets. If you’re running a local model, use its HTTP server mode and interact with it the same way. This keeps process management out of your application and makes retries, timeouts, and streaming much easier to handle.
That makes sense for the model call, but I’m specifically hoping for something with a CLI-like conversational loop that can also invoke actions on my server.
Libraries such as LangChain or a smaller async agent framework can provide the REPL and tool-calling loop, but you don’t strictly need one. A basic implementation is just a conversation history, a loop that sends messages to the model, handling for tool-call responses, and a registry of approved Python functions. Start with that if you only have a few tools, then add a framework when you need memory, streaming, provider switching, or more complex workflows.

The schema validation is the important part. It ensures the model returns structured tool arguments instead of leaving your code to parse free-form text.