consul

Commit Graph

Author	SHA1	Message	Date
Matt Keeler	38f5ddce2a	Add per-agent reconnect timeouts (#8781 ) This allows for client agent to be run in a more stateless manner where they may be abruptly terminated and not expected to come back. If advertising a per-agent reconnect timeout using the advertise_reconnect_timeout configuration when that agent leaves, other agents will wait only that amount of time for the agent to come back before reaping it. This has the advantageous side effect of causing servers to deregister the node/services/checks for that agent sooner than if the global reconnect_timeout was used.	2020-10-08 15:02:19 -04:00
Hans Hasselberg	a932aafc91	add primary keys to list keyring (#8522 ) During gossip encryption key rotation it would be nice to be able to see if all nodes are using the same key. This PR adds another field to the json response from `GET v1/operator/keyring` which lists the primary keys in use per dc. That way an operator can tell when a key was successfully setup as primary key. Based on https://github.com/hashicorp/serf/pull/611 to add primary key to list keyring output: ```json [ { "WAN": true, "Datacenter": "dc2", "Segment": "", "Keys": { "0OuM4oC3Os18OblWiBbZUaHA7Hk+tNs/6nhNYtaNduM=": 6, "SINm887hKTzmMWeBNKTJReaTLX3mBEJKriDyt88Ad+g=": 6 }, "PrimaryKeys": { "SINm887hKTzmMWeBNKTJReaTLX3mBEJKriDyt88Ad+g=": 6 }, "NumNodes": 6 }, { "WAN": false, "Datacenter": "dc2", "Segment": "", "Keys": { "0OuM4oC3Os18OblWiBbZUaHA7Hk+tNs/6nhNYtaNduM=": 8, "SINm887hKTzmMWeBNKTJReaTLX3mBEJKriDyt88Ad+g=": 8 }, "PrimaryKeys": { "SINm887hKTzmMWeBNKTJReaTLX3mBEJKriDyt88Ad+g=": 8 }, "NumNodes": 8 }, { "WAN": false, "Datacenter": "dc1", "Segment": "", "Keys": { "0OuM4oC3Os18OblWiBbZUaHA7Hk+tNs/6nhNYtaNduM=": 3, "SINm887hKTzmMWeBNKTJReaTLX3mBEJKriDyt88Ad+g=": 8 }, "PrimaryKeys": { "SINm887hKTzmMWeBNKTJReaTLX3mBEJKriDyt88Ad+g=": 8 }, "NumNodes": 8 } ] ``` I intentionally did not change the CLI output because I didn't find a good way of displaying this information. There are a couple of options that we could implement later: * add a flag to show the primary keys * add a flag to show json output Fixes #3393.	2020-08-18 09:50:24 +02:00
R.B. Boyer	e853368c23	gossip: Avoid issue where two unique leave events for the same node could lead to infinite rebroadcast storms (#8343 ) bump serf to v0.9.3 to include fix for https://github.com/hashicorp/serf/pull/606	2020-07-21 15:48:10 -05:00
Hans Hasselberg	c4093c87cc	agent: don't let left nodes hold onto their node-id (#7747 )	2020-05-04 18:39:08 +02:00
R.B. Boyer	6adad71125	wan federation via mesh gateways (#6884 ) This is like a Möbius strip of code due to the fact that low-level components (serf/memberlist) are connected to high-level components (the catalog and mesh-gateways) in a twisty maze of references which make it hard to dive into. With that in mind here's a high level summary of what you'll find in the patch: There are several distinct chunks of code that are affected: * new flags and config options for the server * retry join WAN is slightly different * retry join code is shared to discover primary mesh gateways from secondary datacenters * because retry join logic runs in the agent and the results of that operation for primary mesh gateways are needed in the server there are some methods like `RefreshPrimaryGatewayFallbackAddresses` that must occur at multiple layers of abstraction just to pass the data down to the right layer. * new cache type `FederationStateListMeshGatewaysName` for use in `proxycfg/xds` layers * the function signature for RPC dialing picked up a new required field (the node name of the destination) * several new RPCs for manipulating a FederationState object: `FederationState:{Apply,Get,List,ListMeshGateways}` * 3 read-only internal APIs for debugging use to invoke those RPCs from curl * raft and fsm changes to persist these FederationStates * replication for FederationStates as they are canonically stored in the Primary and replicated to the Secondaries. * a special derivative of anti-entropy that runs in secondaries to snapshot their local mesh gateway `CheckServiceNodes` and sync them into their upstream FederationState in the primary (this works in conjunction with the replication to distribute addresses for all mesh gateways in all DCs to all other DCs) * a "gateway locator" convenience object to make use of this data to choose the addresses of gateways to use for any given RPC or gossip operation to a remote DC. This gets data from the "retry join" logic in the agent and also directly calls into the FSM. * RPC (`:8300`) on the server sniffs the first byte of a new connection to determine if it's actually doing native TLS. If so it checks the ALPN header for protocol determination (just like how the existing system uses the type-byte marker). * 2 new kinds of protocols are exclusively decoded via this native TLS mechanism: one for ferrying "packet" operations (udp-like) from the gossip layer and one for "stream" operations (tcp-like). The packet operations re-use sockets (using length-prefixing) to cut down on TLS re-negotiation overhead. * the server instances specially wrap the `memberlist.NetTransport` when running with gateway federation enabled (in a `wanfed.Transport`). The general gist is that if it tries to dial a node in the SAME datacenter (deduced by looking at the suffix of the node name) there is no change. If dialing a DIFFERENT datacenter it is wrapped up in a TLS+ALPN blob and sent through some mesh gateways to eventually end up in a server's :8300 port. * a new flag when launching a mesh gateway via `consul connect envoy` to indicate that the servers are to be exposed. This sets a special service meta when registering the gateway into the catalog. * `proxycfg/xds` notice this metadata blob to activate additional watches for the FederationState objects as well as the location of all of the consul servers in that datacenter. * `xds:` if the extra metadata is in place additional clusters are defined in a DC to bulk sink all traffic to another DC's gateways. For the current datacenter we listen on a wildcard name (`server.<dc>.consul`) that load balances all servers as well as one mini-cluster per node (`<node>.server.<dc>.consul`) * the `consul tls cert create` command got a new flag (`-node`) to help create an additional SAN in certs that can be used with this flavor of federation.	2020-03-09 15:59:02 -05:00
Sarah Christoff	5e26971864	Prune Unhealthy Agents (#6571 ) * Add -prune flag to ForceLeave	2019-10-04 16:10:02 -05:00
Jeff Mitchell	28c84845bf	Update vendoring from go mod. (#5566 )	2019-03-26 17:50:42 -04:00
Matt Keeler	8f0d622a54	Revendor serf to pull in keyring list truncation changes. (#5251 )	2019-01-22 16:07:04 -05:00
R.B. Boyer	b96391ecff	update github.com/hashicorp/{serf,memberlist,go-sockaddr} (#5189 ) This activates large-cluster improvements in the gossip layer from https://github.com/hashicorp/memberlist/pull/167	2019-01-07 15:00:47 -06:00
Paul Banks	9ce10769ce	Update Serf and memberlist (#4511 ) This includes fixes that improve gossip scalability on very large (> 10k node) clusters. The Serf changes: - take snapshot disk IO out of the critical path for handling messages hashicorp/serf#524 - make snapshot compaction much less aggressive - the old fixed threshold caused snapshots to be constantly compacted (synchronously with request handling) on clusters larger than about 2000 nodes! hashicorp/serf#525 Memberlist changes: - prioritize handling alive messages over suspect/dead to improve stability, and handle queue in LIFO order to avoid acting on info that 's already stale in the queue by the time we handle it. hashicorp/memberlist#159 - limit the number of concurrent pushPull requests being handled at once to 128. In one test scenario with 10s of thousands of servers we saw channel and lock blocking cause over 3000 pushPulls at once which ballooned the memory of the server because each push pull contained a de-serialised list of all known 10k+ nodes and their tags for a total of about 60 million objects and 7GB of memory stuck. While the rest of the fixes here should prevent the same root cause from blocking in the same way, this prevents any other bug or source of contention from allowing pushPull messages to stack up and eat resources. hashicorp/memberlist#158	2018-08-09 13:16:13 -04:00
Preetha Appan	fff532cf84	Update serf to pick up clean leave fix	2018-05-04 15:51:55 -05:00
James Phillips	5800474f02	Updates Serf to pickup fix for spammy zero RTT log messages. Fixes #3789.	2018-01-19 14:47:12 -08:00
James Phillips	bcc9aea18f	Updates Serf to pull in new queue depth controls.	2017-12-06 17:06:08 -08:00
Preetha Appan	1af51560d0	Update serf library to pick up coordinate persistence fix	2017-10-21 21:19:43 -05:00
Preetha Appan	f94ba25b9d	REbase master serf	2017-10-20 10:33:59 -05:00
Preetha Appan	9449a60fae	Vendor update serf to pick up fix for out of range ping periods in coordinate subsystem	2017-10-20 10:14:15 -05:00
Frank Schroeder	2567a94a81	serf: monkey patch https://github.com/hashicorp/serf/pull/486	2017-09-26 13:40:26 +02:00
Preetha Appan	c9d5e17410	Update serf to pick up fixes for fsyncing snapshots and panic when coordinates are disabled	2017-08-17 16:35:06 -05:00
Frank Schroeder	8ab88976cf	serf: monkey patch data race in github.com/hashicorp/serf https://github.com/hashicorp/serf/pull/476 This should be replaced when the patch is merged upstream and the library is upgraded.	2017-07-07 09:22:34 +02:00
Preetha Appan	6d172b7059	Update serf to pull in disk space failure recovery changes	2017-07-06 08:58:42 -05:00
Kyle Havlovitz	5d99ee80ca	Add an option to disable keyring file (#3145 ) Also disables keyring file in dev mode.	2017-06-15 15:24:04 -07:00
James Phillips	1021a62ddc	Pulls in Serf logger fix.	2017-05-25 22:02:09 -07:00
James Phillips	9aba84eb1e	Updates Serf to pick up small fixes and coordinate NaN/Inf defenses.	2017-05-25 16:16:37 -07:00
James Phillips	bc2a45a9de	Updates memberlist and Serf (and adds new dependencies). This gets a number of improvements: * Fixed a missing case where gossip would stop flowing to dead nodes for a short while. * Uses the go-sockaddr library to look for private IP addresses, which prefers non-loopback private addresses over loopback ones when trying to automatically determine the advertise address. * Properly seeds Go's random number generator using the seed library. * Serf snapshots no longer have the executable bit set on thie file.	2017-02-08 13:56:07 -08:00
James Phillips	0797d4f92b	Updates Serf library to get relay fixes. https://github.com/hashicorp/serf/pull/447	2017-02-06 08:57:36 -08:00
Kyle Havlovitz	ed5bf524ba	Update serf deps	2017-02-01 20:57:37 -05:00
James Phillips	8ec559a235	Updates Serf to pick up intent queue fix. This fixes #1062 by storing intents per-node instead of in a small, fixed- size circular buffer.	2016-08-08 18:58:44 -07:00
James Phillips	b5f1c41aef	Updates serf to get health score in stats.	2016-06-08 17:20:12 -07:00
James Phillips	a8bafda8b6	Pulls in latest Serf to get flap metric.	2016-06-07 09:22:41 -07:00
Sean Chittenden	bf8c860663	Update Serf to include `serf.NumNodes()`	2016-03-23 22:10:50 -07:00
Sean Chittenden	25b09713b9	Manage dependencies via Godep Embrace the future and use Go 1.6's vendor support via Godep. Go 1.5 users should `export GO15VENDOREXPERIMENT=1`	2016-02-12 16:50:37 -08:00

31 Commits