On this page
Why this post exists
Elasticsearch ships with two famously sharp footguns: an HTTP API that does whatever you ask if you can reach it, and a configuration surface that splits security across half a dozen files. Our monitor-agent catches the operational signals (cluster RED, disk watermark, GC, shard skew) the moment they happen. The list below is everything you have to fix yourself — once — before those alerts even matter.
Every command below was run on a fresh `elasticsearch:8.11.0` container on 2026-05-26 — exact request/response captures live in [the spec](https://github.com/elk-utilities/homelab-apps/blob/main/.ai/specs/security-101-best-practices-blog.md). The fixes work on 7.x with caveats called out in the version notes.
Is xpack.security actually on?
On 8.x, security defaults to ON if you initialize a new cluster with the bundled keystore. Anyone who set `xpack.security.enabled: false` to `make it work` during a deploy has an open HTTP API at port 9200. There is no metric for this — the cluster looks healthy from the outside.
Check: probe the security feature flag
curl -s "$ES_URL/_xpack" \
| jq '.features.security.enabled'What to look for: `false` means anybody with a route to port 9200 can read, write, and execute scripts against your cluster. On 8.x default this should be `true`.
Check: can you query without credentials?
The definitive test. If this returns `200`, security is off, period.
curl -s -o /dev/null -w 'HTTP %{http_code}\n' \
"$ES_URL/_cluster/health"What to look for: `HTTP 401` means security is on. `HTTP 200` means anyone can read your cluster.
Fix: enable security on a running cluster
On 8.x with auto-config the keystore already has a built-in `elastic` superuser password — `elasticsearch-reset-password -u elastic` prints it. On 7.x or any cluster that started with security disabled, the flag flip below plus a rolling restart is the path.
Verify it worked: # After the rolling restart, set the elastic-user password and verify
# auth is enforced:
bin/elasticsearch-reset-password -u elastic
curl -s -o /dev/null -w 'HTTP %{http_code}\n' "$ES_URL/_cluster/health"
# expect: HTTP 401
curl -s -o /dev/null -w 'HTTP %{http_code}\n' -u elastic:'<password>' "$ES_URL/_cluster/health"
# expect: HTTP 200
Note: If you cannot do a rolling restart immediately, restrict network access at the firewall / NetworkPolicy / VPC security group level *first*. The goal is to remove unauthenticated reachability — security on the cluster itself is the durable fix.
Are you on an EOL major?
Elasticsearch 5.x and 6.x reached end-of-life in 2019 and 2022 respectively. 7.x stopped receiving security backports in August 2023. Clusters on these majors are, in practical terms, indefinitely exposed.
Check: what version are you on?
What to look for: Anything but `OK:` is something a customer audit will flag.
Fix: plan a real migration, not an in-place upgrade
Across two majors (6→8, 7→9 once it ships) an in-place upgrade is not supported — `elasticsearch-shard remove-corrupted-shard` is not a migration plan. The supported path is reindex-from-remote into a new cluster.
Verify it worked: # Poll the task:
curl -s -u elastic:'<password>' \
"$NEW_ES_URL/_tasks?actions=*reindex&detailed=true" \
| jq '.nodes[].tasks[] | {action,description,running_time_in_nanos}'
Note: Allowlist the source cluster on the destination via `reindex.remote.whitelist: "OLD-CLUSTER:9200"` in `elasticsearch.yml` before running the request, or you'll see `not in the reindex.remote.whitelist`.
Are you a single-node cluster pretending to be HA?
One data node has zero redundancy. The cluster appears green only because replicas haven't been provisioned anywhere they could land — the moment you create a replica index, you go YELLOW and stay there. The next disk failure is a full outage.
Check: count the data nodes
curl -s "$ES_URL/_cluster/health" \
| jq '{status, data_nodes: .number_of_data_nodes}'What to look for: `{"status":"green","data_nodes":1}` is the classic "single-node pretending to be HA" pattern.
Fix: scale to 3 nodes (minimum quorum)
Two data nodes is technically HA but produces brain-split risk for the master quorum. Three is the canonical minimum. The exact provisioning command depends on your platform — the snippet below is the Helm chart from the Elastic operator, with the operator-friendly affinity rule that prevents two pods landing on the same physical node.
Verify it worked: curl -s -u elastic:'<password>' "$ES_URL/_cat/nodes?v" \
| awk 'NR>1 {print $1, $10}' | sort -u
# Expect three distinct host IPs.
Did somebody disable shard allocation and forget to turn it back on?
`cluster.routing.allocation.enable = none` is the canonical "start of a rolling restart" setting. It's also the canonical "got distracted by an incident and forgot to re-enable" setting. A cluster in this state cannot recover from any node failure or balance new indices.
Check: is allocation enabled?
curl -s "$ES_URL/_cluster/settings?include_defaults=true&filter_path=*.cluster.routing.allocation.enable" \
-u elastic:'<password>' \
| jqWhat to look for: Look at `transient.cluster.routing.allocation.enable` and `persistent.…`. Anything other than `all` (or missing — defaults to `all`) is a finding.
Fix: re-enable allocation
Verify it worked: curl -s -u elastic:'<password>' \
"$ES_URL/_cluster/settings?include_defaults=true&filter_path=*.cluster.routing.allocation.enable" \
| jq '.. | .enable? // empty' | sort -u
# Expect a single line: "all"
Note: If shards still don't move after this, run `GET _cluster/allocation/explain` to see the next blocker (disk watermark, awareness rules, max_retries).
Are there indices with zero replicas?
An index with `number_of_replicas: 0` has no copy of its primary shards. If the node hosting them dies before a snapshot, the data is gone. This is sometimes intentional for ephemeral indices but is almost never the right choice for anything customer-facing.
Check: list indices with rep=0
curl -s -u elastic:'<password>' \
"$ES_URL/_cat/indices?h=index,rep&format=json" \
| jq '.[] | select(.rep == "0") | .index'What to look for: Any index name in the output is at data-loss risk on the next disk failure.
Fix: set replicas to at least 1
curl -X PUT "$ES_URL/<index-name>/_settings" \
-H 'Content-Type: application/json' \
-u elastic:'<password>' \
-d '{"index": {"number_of_replicas": 1}}'Verify it worked: curl -s -u elastic:'<password>' \
"$ES_URL/_cat/indices/<index-name>?h=index,rep"
# Expect: <index-name> 1
Note: On a single-node cluster this will move the cluster to YELLOW — there's nowhere for the new replica to land. That YELLOW is informative: it tells you the topology was always at risk. See section 4 to fix the underlying single-node problem.