Why the Clash External Controller API matters for node failover
A manual node switch is acceptable when you are watching the Clash dashboard, but it is a poor recovery strategy for a workstation, home gateway, test runner, or small server that must remain reachable while you are away. A proxy can become unusable for several different reasons: the upstream server may be overloaded, a transit route may lose packets, DNS may resolve to an unreachable address, or the local client may still show a connected tunnel even though new HTTPS handshakes fail. In each case, clicking a different node works only after somebody notices the failure.
The Clash External Controller API exposes the information and controls required to replace that manual loop with a small automation layer. A script can query active connections, inspect proxy groups, measure a known URL through a selected policy, and send a request that changes the currently selected node. Modern Mihomo-based clients generally expose compatible controller endpoints, although menu names, default ports, and permission dialogs differ between Clash Verge Rev, Mihomo Party, Clash for Android, and other clients. The API is not a magic optimizer: it executes the policy your script defines, so reliable automation begins with conservative measurements and explicit failure boundaries.
The most important design decision is to separate observation from selection. First collect enough evidence to decide that a route is unhealthy. Then select a replacement from a known group. Finally verify that the replacement actually carries traffic. Switching after one slow response creates unnecessary churn, while never switching because the control endpoint itself is unreachable leaves the system stuck. A practical controller therefore uses a probe URL, a latency threshold, a failure counter, a cooldown period, and a final verification request.
This guide assumes that you administer the Clash instance and have permission to alter its routing state. The examples are intended for personal devices, laboratories, and networks where automated proxy changes are allowed. Do not use an external controller token from a shared chat, expose the controller port to the public internet, or automate changes that violate an employer, school, provider, or local network policy.
Prepare the controller endpoint and proxy group
Before writing a script, open the client settings and identify the External Controller address. A common local value is 127.0.0.1:9090, but your client may use another port or bind to an IPv6 loopback address. The controller address is independent from the mixed port or SOCKS port used by applications. Sending API requests to the traffic port will produce confusing errors because a proxy listener does not understand controller routes.
Recent Clash and Mihomo configurations commonly protect the API with an authentication token. A minimal local configuration can look like this:
external-controller: 127.0.0.1:9090
secret: replace-with-a-long-random-token
proxy-groups:
- name: AUTO_FAILOVER
type: fallback
url: https://www.gstatic.com/generate_204
interval: 300
proxies:
- Node-A
- Node-B
- Node-C
The fallback group is a useful starting point because it already expresses ordered active/passive recovery. A script can still select a specific child when the built-in probe is not sensitive enough for your workload. If you prefer continuous latency ranking, use a url-test group; if you need predictable manual control, use a select group and let the script make the decision. Do not confuse these modes. A selector preserves explicit intent, a fallback group prefers the first healthy member in its list, and a URL test group attempts to promote the member with the best measured result.
Group names and node names are case-sensitive in practical API use. The visible label may also contain spaces, emoji, regional suffixes, or characters that require URL encoding. Always read the group list from the API instead of assuming that a provider’s subscription names remain unchanged after a refresh. A renamed node is a normal operational event, not necessarily evidence that the controller has failed.
You can verify the controller manually with an authenticated request:
curl -sS \
-H "Authorization: Bearer replace-with-a-long-random-token" \
http://127.0.0.1:9090/proxies
A successful response should contain a JSON object with proxy groups and their current selections. If you receive 401 Unauthorized, check the token and the exact Bearer spelling. A connection refusal usually means the client is closed, the controller is disabled, or the port is different. A timeout may indicate that the endpoint is bound only to another address or that a local firewall is filtering the request. Solve this layer before investigating nodes; a script cannot perform failover when it cannot reach the control plane.
9090 through a router without an authenticated, encrypted management path, and do not place the secret directly in a public repository.
Build a guarded node-switching script
The controller route for a proxy group is commonly a GET request to inspect the group and a PUT request to change its selected proxy. The exact endpoint is usually shaped like /proxies/{group}, with the group name URL-encoded. A robust script should not blindly issue a PUT. It should confirm that the group exists, ensure the requested node is a member, avoid selecting the node that is already active, and check the response status before declaring success.
The following Python example uses only the standard library. It measures a public HTTP endpoint through a local proxy port, records consecutive failures, applies a cooldown, and changes the AUTO_FAILOVER group when the active node fails repeatedly. Set the local traffic port to the port used by your Clash profile, not the external controller port.
#!/usr/bin/env python3
import json
import os
import time
import urllib.error
import urllib.parse
import urllib.request
CONTROLLER = os.getenv("CLASH_CONTROLLER", "http://127.0.0.1:9090")
SECRET = os.environ["CLASH_SECRET"]
GROUP = os.getenv("CLASH_GROUP", "AUTO_FAILOVER")
PROXY_PORT = int(os.getenv("CLASH_MIXED_PORT", "7890"))
PROBE_URL = os.getenv("CLASH_PROBE_URL", "https://www.gstatic.com/generate_204")
TIMEOUT = float(os.getenv("CLASH_TIMEOUT", "8"))
THRESHOLD_MS = float(os.getenv("CLASH_THRESHOLD_MS", "1800"))
MAX_FAILURES = 3
COOLDOWN = 300
last_switch = 0.0
failures = 0
def api(path, method="GET", payload=None):
body = None
headers = {"Authorization": f"Bearer {SECRET}"}
if payload is not None:
body = json.dumps(payload).encode()
headers["Content-Type"] = "application/json"
request = urllib.request.Request(
CONTROLLER.rstrip("/") + path,
data=body,
headers=headers,
method=method,
)
with urllib.request.urlopen(request, timeout=TIMEOUT) as response:
return response.status, json.loads(response.read().decode())
def probe():
proxy = urllib.request.ProxyHandler({
"http": f"http://127.0.0.1:{PROXY_PORT}",
"https": f"http://127.0.0.1:{PROXY_PORT}",
})
opener = urllib.request.build_opener(proxy)
started = time.monotonic()
request = urllib.request.Request(PROBE_URL, method="GET")
with opener.open(request, timeout=TIMEOUT) as response:
response.read(128)
return (time.monotonic() - started) * 1000
def main():
global failures, last_switch
try:
latency = probe()
print(f"probe={latency:.0f}ms")
if latency <= THRESHOLD_MS:
failures = 0
return
failures += 1
except (urllib.error.URLError, TimeoutError, OSError) as error:
failures += 1
print(f"probe failed: {error}")
if failures < MAX_FAILURES:
print(f"failure_count={failures}; waiting for confirmation")
return
if time.time() - last_switch < COOLDOWN:
print("cooldown active; no switch")
return
encoded_group = urllib.parse.quote(GROUP, safe="")
_, group = api(f"/proxies/{encoded_group}")
current = group["now"]
candidates = [name for name in group.get("all", []) if name != current]
if not candidates:
raise RuntimeError("group has no alternate members")
target = candidates[0]
api(f"/proxies/{encoded_group}", "PUT", {"name": target})
last_switch = time.time()
failures = 0
print(f"switched {current} -> {target}")
if __name__ == "__main__":
main()
This example deliberately keeps the policy simple. It chooses the first alternate member rather than pretending that a single probe can rank every node accurately. In production, you can retrieve each candidate, temporarily test it, and score it using several samples. Another option is to let a Mihomo url-test group perform the ranking and use the API only to read the winner. The safest approach depends on whether your goal is availability, latency, geographic consistency, or minimizing the number of control changes.
The failure counter is essential. Networks produce isolated packet loss, captive portal redirects, transient DNS errors, and slow responses during radio handoffs. Switching after every exception can create a loop in which the script moves away from a perfectly serviceable node and then returns to it on the next run. Requiring three failures, spacing checks with a scheduler, and adding a cooldown makes the decision less reactive. For streaming applications, include a longer application-level check because a fast HTTP status does not prove that a WebSocket or long response stream will remain stable.
Schedule checks without creating a control loop
A scheduled job should run frequently enough to reduce downtime but slowly enough to avoid turning the probe endpoint into a source of noise. A five-minute interval is reasonable for a home gateway, while a laptop that changes networks may need a check after wake, resume, or Wi-Fi association. Run the script under the same user that can reach the local proxy and store the secret in an environment file with restrictive permissions rather than embedding it in the command line.
On Linux, a systemd timer provides clearer logs and dependency control than an endlessly running shell loop:
# /etc/systemd/system/clash-failover.service
[Unit]
Description=Clash node failover check
[Service]
Type=oneshot
User=clash
EnvironmentFile=/etc/clash/failover.env
ExecStart=/usr/local/bin/clash-failover.py
# /etc/systemd/system/clash-failover.timer
[Unit]
Description=Run Clash failover check every five minutes
[Timer]
OnBootSec=90s
OnUnitActiveSec=5min
Persistent=true
[Install]
WantedBy=timers.target
After creating the units, reload systemd, enable the timer, and inspect its journal. The important operational question is not only whether a switch occurred, but why it occurred. Log the timestamp, group, previous node, candidate node, probe target, latency, and error category. Avoid logging the controller token, subscription URL, authorization headers, or full response bodies when they may contain sensitive metadata. A short record such as “three HTTPS probe timeouts; switched from Node-A to Node-B” is more useful than a credential-bearing debug dump.
Windows users can create a Task Scheduler task that runs Python with an environment file or a wrapper script. Choose “Run whether user is logged on or not” only when the machine’s security model permits it, and make sure the task does not start several copies concurrently. macOS users can use a LaunchAgent for a per-user setup or a LaunchDaemon for a managed service. In both systems, confirm that the task starts after Clash and its controller listener are ready; otherwise the first run may report a false outage.
Add a lock if the scheduler can overlap jobs. A slow probe followed by a second invocation can produce two independent decisions and two rapid node changes. A lock file, operating-system mutex, or short-lived database lease prevents concurrent execution. Also define what happens when the controller is unavailable. The correct behavior is usually “record the control-plane failure and leave the current configuration unchanged,” not “rewrite the YAML from a half-complete script.” Configuration rewrites can discard provider updates, manual selections, or client-managed fields.
Troubleshoot false switches, stuck nodes, and API errors
Start troubleshooting by identifying which layer failed. If the API request returns a valid group but the probe fails, the control plane is healthy and the traffic path needs investigation. If the probe succeeds but the PUT request returns an error, inspect group permissions, URL encoding, and the selected node name. If the PUT succeeds yet traffic remains broken, verify that applications actually use the mixed port or TUN interface associated with the Clash instance being controlled. It is common to operate two clients at once and send API commands to one while the browser uses the other.
A 404 response often means that the endpoint shape does not match the core version or that the group name was not encoded correctly. A group called Travel / Auto must not be inserted into a URL as raw text; encode it before constructing the path. A 400 response usually indicates malformed JSON or a node name that is not accepted by the group. A 401 points to the token, while a 403 may indicate a client-specific permission rule. Capture the status code and a sanitized response message, then compare the request with the API documentation shipped by your core or client.
False latency alarms are usually measurement problems. A probe may be served from a nearby cache, respond before the application’s real API is reachable, or succeed through a different DNS family. Use a stable HTTPS endpoint and keep the target consistent across comparisons. If your workload depends on a particular vendor API, add a second permitted health check that resembles the real connection without sending private data. Measure several samples, discard obvious local scheduling pauses, and consider packet loss as well as elapsed time. A node with 120 milliseconds of latency and five percent loss may be worse than one with 250 milliseconds and no loss.
DNS deserves separate attention. With fake-IP mode, the application may resolve a synthetic address locally while the core resolves the actual destination later. A script that probes a hostname through the operating system resolver may therefore test a different path from the application. Likewise, a successful controller request to 127.0.0.1 says nothing about whether remote DNS, rule providers, or the selected proxy can resolve the target. Compare Clash logs, DNS mode, rule matches, and connection records rather than concluding that “the node is dead” from one application error.
Protect against oscillation with hysteresis. Require the replacement to remain healthy for more than one sample before switching back, and use a recovery threshold that is better than the failure threshold. For example, switch away after three failures above 1,800 milliseconds, but switch back only after three consecutive samples below 700 milliseconds. Keep a maximum switch count per hour and alert when it is exceeded. Repeated changes can indicate provider instability, a broken probe, an overloaded local machine, or a subscription that contains duplicate endpoints.
Operational rule: automatic failover should reduce human intervention, not remove human visibility. Keep a manual selector available, document the command that disables the scheduled job, and review switch logs after changing providers, rules, DNS mode, or TUN settings.
A safer production pattern for advanced operators
For a dependable deployment, place the automation behind a small configuration file that defines the controller address, group name, probe policy, thresholds, and candidate strategy. Validate every value before the first API call. Restrict the group to nodes that are genuinely interchangeable; mixing a low-latency regional node, a high-latency backup, and a special-purpose egress in one automatic pool can create surprising application behavior. Keep sensitive services on a deliberate selector group when stable identity matters more than recovery speed.
Separate node health from service health. A proxy can pass a generic connectivity test while a particular API is blocked, rate-limited, or returning application errors. Conversely, a vendor endpoint can be temporarily degraded while the node remains useful for ordinary browsing. Use clear names such as GENERAL_AUTO, STREAMING_MANUAL, and WORK_API_SELECT rather than one universal group that every rule depends on. This makes both automation and incident review easier.
Test failure scenarios before trusting the scheduler. Temporarily stop the selected client, block the probe route in a controlled lab, use an intentionally invalid controller token, and rename a candidate node in a disposable profile. Confirm that the script logs the correct diagnosis, avoids leaking credentials, respects cooldowns, and leaves the configuration intact when no safe candidate exists. Then test recovery by restoring the route and confirming that the script can observe the new state without endlessly switching.
Compared with GUI-only switching, scripts bundled into provider dashboards often expose fewer diagnostic details and can hide whether a change affected a selector, fallback group, or TUN route; simple shell loops, meanwhile, commonly lack token protection, URL encoding, cooldowns, and rollback awareness. Clash V.CORE gives you a clearer foundation for this workflow: an accessible controller model, inspectable groups, configurable routing, and a practical interface for validating behavior before automation takes over. If you want to build the setup on a supported client and keep the management surface under your control, download Clash V.CORE and begin with a local, authenticated controller before enabling scheduled failover.
// Editor's Pick
Clash V.CORE for reliable node automation
Build safer API-driven switching with visible groups, practical health checks, and a control workflow that remains understandable when a node fails.
- Inspect active groups and selected nodes
- Support authenticated controller workflows
- Combine fallback and URL-test strategies
- Validate proxy behavior before switching
- Keep manual control during incidents