Why the External Controller matters for automated switching

Clash is often introduced as a desktop application: import a profile, choose a proxy group, and let the rule engine decide where each connection goes. That workflow is sufficient for casual browsing, but it becomes limiting when the proxy must react to measurable conditions. A development machine may need to move away from a slow node before a package installation begins; a build server may need to avoid a provider whose subscription has just returned errors; a test runner may need to verify several exits without a person clicking through a GUI. This is where the Clash External Controller API becomes useful.

The controller is an HTTP interface exposed by the Clash or Mihomo core. It gives local tools a way to inspect running state, read proxy groups, query traffic and connection information, test node delay, and change the active member of a selector group. The API does not replace YAML rules. Instead, it adds an operational layer above the configuration: YAML describes the available groups and policies, while a script can observe current conditions and make a controlled decision at runtime.

That distinction is important. An API request can select a different member inside a group, but it cannot magically repair a malformed provider, create a missing policy group, or make an incompatible client support every Mihomo endpoint. Clash Verge, Clash Verge Rev, Clash for Windows forks, ClashX variants, Clash for Android, and Mihomo-based clients may expose slightly different menus, ports, and endpoint behavior. Always identify the actual core and verify its API documentation before copying a command from another installation.

ℹ Scope: The examples assume a controller bound to a trusted local interface. Use automated switching only for networks, accounts, and services you are authorized to operate, and follow provider, employer, and local network policies.

Controller basics: address, secret, and API discovery

The two settings that matter most are the controller address and its secret. In a typical configuration, external-controller contains an address such as 127.0.0.1:9090, while secret protects requests with an authorization token. A local-only binding is the safest starting point because it prevents other machines on the LAN from reaching the management interface. If you expose the controller on 0.0.0.0, every reachable interface becomes part of your attack surface unless a firewall or a private management network restricts access.

external-controller: 127.0.0.1:9090
secret: change-this-to-a-long-random-value

After changing the profile, restart or reload the core according to the client’s workflow. Do not assume that editing a YAML file changes the already running process. Some clients generate a temporary profile, some merge several files, and some keep controller settings in a separate preferences page. Check the effective configuration shown by the application, then test the endpoint from the same machine.

curl -fsS \
  -H 'Authorization: Bearer change-this-to-a-long-random-value' \
  http://127.0.0.1:9090/version

A successful response normally identifies the running core and version. This first request is valuable because it separates an authentication problem from a routing or node problem. A connection refusal usually means the controller is disabled, bound to another port, or blocked by a local process conflict. A 401 or 403 response points toward a missing or incorrect token. A valid JSON response confirms that later failures should be investigated at the endpoint, group, provider, or node layer rather than at the controller socket.

Keep the token out of shell history, shared screenshots, CI logs, and public repositories. For a quick local test, an environment variable is preferable to placing the secret directly in a script. On a server, store it in a protected secret manager or a file readable only by the service account. The controller is not an ordinary browsing API: whoever can call its mutation endpoints may be able to redirect traffic, inspect connections, or disrupt an entire machine’s network policy.

The endpoints used most often

Purpose Typical endpoint Operational use
Identify the core GET /version Confirm that the controller is reachable and learn the core version.
List groups GET /proxies Discover selector, url-test, fallback, and load-balance groups.
Inspect one group GET /proxies/{name} Read the current member and available choices.
Test delay GET /proxies/{name}/delay Measure a node or group against a probe URL.
Switch a selector PUT /proxies/{name} Set the active member of a compatible proxy group.
Read providers GET /providers/proxies Inspect provider health and refresh state when supported.

Endpoint names and response fields can vary between Clash forks and Mihomo releases, so treat this table as a working map rather than a promise of identical behavior. Names containing spaces or non-ASCII characters must be URL-encoded. A group called Auto Select should not be inserted into a URL as a raw, unescaped string. In scripts, let the HTTP library encode path components instead of concatenating untrusted names manually.

Inspect groups before changing anything

The safest automation begins with observation. Query the full proxy collection and locate the group you intend to control. A common mistake is to switch a provider name or an individual node when the traffic rules actually point to a higher-level selector. If rules send traffic to Proxy, changing a child group named US Nodes may have no visible effect unless Proxy already selects that child group. The API can report the current topology, but it cannot infer your intended policy.

curl -fsS \
  -H "Authorization: Bearer $CLASH_SECRET" \
  http://127.0.0.1:9090/proxies | jq '.proxies.Proxy, .proxies["US Nodes"]'

Read at least the group type, current selection, and available members. A selector generally accepts an explicit mutation request. A url-test group may calculate its own choice and reject or ignore manual selection depending on the core. A fallback group is governed by health ordering, while a load-balance group may use hashing or round-robin behavior rather than one permanent active node. Automation should therefore target a deliberately created selector when deterministic control is required.

Give important groups stable names. Avoid names that change every time a provider refreshes, such as labels generated from timestamps or temporary import metadata. Stable group names make scripts easier to audit and allow a profile to evolve internally without forcing every deployment tool to change. It is also useful to separate concerns: create one group for ordinary browsing, another for development services, and a third for a narrowly defined test workflow. Switching a general-purpose group from a build script can otherwise surprise users who are browsing at the same time.

Operational rule: Never select the first item returned by /proxies merely because it is convenient. Filter candidates by the exact group, exclude DIRECT or REJECT when appropriate, preserve the current choice until a replacement passes your checks, and record the reason for every change.

Before mutation, capture the current state. A small JSON record containing the group name, old member, candidate, measured delay, timestamp, and script version is enough to explain a later incident. This also makes rollback simple. If a deployment temporarily changes Build Proxy, the cleanup step can restore the original member instead of guessing which node a human had selected earlier.

Measure node latency without confusing speed with health

The delay endpoint is useful, but latency is only one signal. A node can answer a lightweight probe quickly while failing long-lived TLS connections, streaming responses, large downloads, or a particular destination. Conversely, a geographically distant node may show a higher round-trip time while providing a more stable path for the service that matters. Choose a probe URL that is permitted in your environment and representative of the workload, then compare candidates under the same timeout and URL conditions.

curl -fsS --get \
  -H "Authorization: Bearer $CLASH_SECRET" \
  --data-urlencode 'url=https://www.gstatic.com/generate_204' \
  --data-urlencode 'timeout=5000' \
  "http://127.0.0.1:9090/proxies/Node%20A/delay"

Do not treat one successful response as proof of reliability. Run several samples, discard obvious outliers, and use a median or a percentile rather than blindly selecting the smallest number. A practical script might require three successful probes within a short window, reject any result above a threshold, and prefer the current node when the difference is insignificant. This hysteresis prevents constant switching when two exits alternate between 180 and 190 milliseconds.

Probe failures should be classified. A timeout can indicate a dead node, a blocked probe destination, a DNS issue, or a controller request that was incorrectly encoded. A successful API response with a very high delay is different from an HTTP error returned by the probe target. Store the status and error body where possible, but redact tokens and any sensitive URLs before sending logs to a central system.

The probe URL should also match the rule path you are trying to evaluate. If a node is being selected for a private registry, probing a public landing page only tests a broad approximation. When permitted, use an endpoint that returns a small deterministic response from the same service family. Do not over-poll provider nodes: frequent tests consume bandwidth, create noisy metrics, and can trigger provider rate limits. A scheduler running every few minutes is usually more responsible than a loop that probes every second.

Build a safe automatic switching workflow

A reliable switching loop has four phases: discover, measure, decide, and mutate. Discovery verifies that the target group and candidate names still exist. Measurement gathers comparable results. Decision applies thresholds, exclusions, and a cooldown. Mutation changes the group only after the replacement has passed validation. Keeping these phases separate makes the script easier to test and prevents a transient API error from being interpreted as permission to switch to an arbitrary member.

The following Python example illustrates the decision logic. It deliberately uses a selector group and receives the controller secret through an environment variable. Adapt the endpoint details to the core installed in your client, and verify the response schema before using it in production.

import os
import statistics
import time
import requests

BASE = "http://127.0.0.1:9090"
GROUP = "Build Proxy"
NODES = ["Node A", "Node B", "Node C"]
TOKEN = os.environ["CLASH_SECRET"]
HEADERS = {"Authorization": f"Bearer {TOKEN}"}
PROBE = "https://www.gstatic.com/generate_204"

def delay(node):
    values = []
    for _ in range(3):
        response = requests.get(
            f"{BASE}/proxies/{node}/delay",
            params={"url": PROBE, "timeout": 5000},
            headers=HEADERS,
            timeout=7,
        )
        response.raise_for_status()
        values.append(int(response.json()["delay"]))
        time.sleep(0.4)
    return statistics.median(values)

state = requests.get(f"{BASE}/proxies/{GROUP}", headers=HEADERS, timeout=5)
state.raise_for_status()
current = state.json()["now"]

results = {}
for node in NODES:
    try:
        results[node] = delay(node)
    except (requests.RequestException, ValueError):
        continue

healthy = {node: ms for node, ms in results.items() if ms < 800}
if healthy:
    best = min(healthy, key=healthy.get)
    if best != current and healthy[best] + 80 < results.get(current, 99999):
        change = requests.put(
            f"{BASE}/proxies/{GROUP}",
            json={"name": best},
            headers=HEADERS,
            timeout=5,
        )
        change.raise_for_status()
        print(f"switched {current} -> {best}: {healthy[best]} ms")

The important part is not the exact threshold. It is the margin between the current node and the candidate. Without a margin, the script may oscillate whenever measurements are close. Add a cooldown so that a successful change cannot be followed by another change for several minutes. In a shared workstation, also check whether a user has manually selected a group recently; an automation job should not silently override deliberate interactive control unless that behavior has been agreed in advance.

For a server deployment, use a service manager or scheduler with clear limits. Set request timeouts, cap the number of retries, and fail closed: if the controller becomes unavailable, do not send a mutation request to a different address. Emit structured logs with the decision, but avoid logging the bearer token, subscription URL, full proxy credentials, or raw connection metadata. If the script runs in CI, protect secrets from command echo and make the job’s permissions as narrow as possible.

Handle failed providers, stale nodes, and rollback

A node switching script is only as good as its provider assumptions. Subscription providers can return temporary HTTP errors, expired data, malformed proxy definitions, or an otherwise valid list with an empty group. Refreshing a provider and switching a node are different operations. Refresh only when the profile and provider contract allow it, and wait for the refresh result before querying candidates again. If the provider is still unavailable, keep the last known-good selection rather than replacing it with DIRECT by accident.

Provider health should be observed separately from node latency. A fast node that disappears after a provider update is not an ordinary timeout; it is a configuration lifecycle event. Record the provider name, refresh timestamp, candidate count, and active group member. When a refresh removes the selected node, choose a replacement from a validated allowlist and notify the operator. For critical services, require an explicit approval step instead of allowing a background job to make an irreversible-looking change.

Build a rollback path before enabling automation. The simplest rollback is an API request that restores the captured member. A stronger design stores the previous state in a short-lived file with restrictive permissions and expires it after the deployment window. If the controller reports that the old member no longer exists, the script should report a manual intervention requirement rather than selecting a similarly named node. Names can look alike while representing different providers, regions, or policy terms.

Test failure cases with a disposable profile. Stop the core, change the secret, remove a candidate, make the probe URL unreachable, and return an invalid group name. Confirm that each case produces a clear error and leaves the active group unchanged. Also verify concurrent behavior: two scheduled jobs may read the same old member and then make conflicting decisions. A lock file, distributed lock, or single-instance service prevents one automation process from undoing another’s work.

Secure the controller on development machines and servers

Keep the controller on loopback whenever the script runs on the same host. If remote administration is necessary, bind it only to a private management address, restrict it with host firewall rules, and place it behind an authenticated encrypted channel such as an approved SSH tunnel or a properly configured reverse proxy. Do not publish the controller directly to the public internet. An exposed controller is effectively a remote network control plane, not a harmless status page.

Use a long random secret and rotate it when a machine changes ownership, a developer leaves a project, logs may have leaked, or a backup is restored to another environment. Review access logs if your deployment includes an intermediary proxy. Avoid putting the secret in a URL because URLs are commonly copied into browser history, monitoring systems, and reverse-proxy logs. Prefer an authorization header and reject requests that do not use the expected method and content type.

Limit automation permissions by architecture even when the core itself does not provide fine-grained roles. A read-only monitoring process can be separated from a switching process. The latter can run only on a trusted host, accept a fixed group allowlist, and refuse names containing unexpected path characters. Validate JSON responses before using fields such as now or all; an error object should never be treated as an empty healthy-node list.

Finally, remember that the controller can reveal operationally sensitive information: active destinations, connection metadata, provider names, and selected nodes. Protect debug exports and redact them before attaching them to issue trackers. On shared development machines, do not assume that every local account is trusted merely because the address is 127.0.0.1; local malware and untrusted desktop extensions can still make requests to loopback services.

Compared with GUI-only clients, which often require repetitive clicks and may hide provider failures behind a compact status icon, or simple shell wrappers that switch nodes without measuring health, the Clash External Controller API offers a consistent path to inspection, latency-aware decisions, cooldowns, audit logs, and rollback. Clash V.CORE brings that controller-centered workflow into a maintained environment where development machines and server deployments can keep policy groups visible while automation handles routine checks. If you want to turn the procedure above into a dependable daily tool, download Clash V.CORE and start with a loopback-only controller before expanding your deployment.

// Editor's Pick

Clash V.CORE for API-driven proxy operations

Build safer node-switching workflows with clear group visibility, practical controller access, and a reliable base for development and server-side automation.

  • Inspect proxy groups before every change
  • Measure node delay with repeatable probes
  • Keep provider failures separate from node health
  • Apply cooldowns and rollback decisions
  • Protect local controller access with a secret
Get Clash V.CORE →