Local inference, familiar workflow.
Orca connects to your local server; it does not install or download models, bundle inference, or manage Ollama and LM Studio processes. Start the server and load your model before connecting.
- Open Orca’s settings and select Ollama, LM Studio, or Local OpenAI-compatible.
- Confirm the endpoint, refresh the model list, or enter a Model ID manually.
- Start chatting. No API key is required if your server does not require one.
Chat first. Agent tools by choice.
Local and editable profiles begin in Chat mode. An isolated two-step function-call probe must pass for the exact provider, origin, endpoint, model, reasoning configuration, and probe version. You must then explicitly enable Agent tools for that binding.
The probe contains no project files, editor context, instructions, skills, task lists, conversation history, or real Orca tool schema. A passing probe confirms one narrow protocol exchange—not reliable planning or tool use on a real project.
The tested qwen2.5-coder:3b and llama-3.2-3b-instruct models worked for Chat but did not pass Orca’s strict Agent probe. More capable tool-use models may be needed.
Know where your requests go.
Orca uses a streamed Chat Completions transport. Prompts, relevant editor context, and project content returned by tools may be sent to your selected provider. A local endpoint can still forward data elsewhere; Orca cannot verify what the server does after receiving a request.
Custom compatible providers may expose incomplete model or capability metadata. See the safety and privacy overview before configuring an endpoint.