consul

Commit Graph

Author	SHA1	Message	Date
Paul Banks	88388d760d	Support Agent Caching for Service Discovery Results (#4541 ) * Add cache types for catalog/services and health/services and basic test that caching works * Support non-blocking cache types with Cache-Control semantics. * Update API docs to include caching info for every endpoint. * Comment updates per PR feedback. * Add note on caching to the 10,000 foot view on the architecture page to make the new data path more clear. * Document prepared query staleness quirk and force all background requests to AllowStale so we can spread service discovery load across servers.	2018-10-10 16:55:34 +01:00
Igal Shprincis	e1fe3af37f	watch: don't set TLSConfig.Address explicitly (#4727 ) Don't set the value of TLSConfig.Address explicitly. This will make sure env vars like CONSUL_TLS_SERVER_NAME are taken into account for the connection. Fixes #4718.	2018-10-08 22:01:17 +02:00
Paul Banks	e8ba527f23	Add a Close method to cache that stops background goroutines. (#4746 ) In a real agent the `cache` instance is alive until the agent shuts down so this is not a real leak in production, however in out test suite, every testAgent that is started and stops leaks goroutines that never get cleaned up which accumulate consuming CPU and memory through subsequent test in the `agent` package which doesn't help our test flakiness. This adds a Close method that doesn't invalidate or clean up the cache, and still allows concurrent blocking queries to run (for up to 10 mins which might still affect tests). But at least it doesn't maintain them forever with background refresh and an expiry watcher routine. It would be nice to cancel any outstanding blocking requests as well when we close but that requires much more invasive surgery right into our RPC protocol since we don't have a way to cancel requests currently. Unscientifically this seems to make tests pass a bit quicker and more reliably locally but I can't really be sure of that!	2018-10-04 11:27:11 +01:00
Paul O'Connor	6b7f03911e	Fix prometheus error message (#4745 )	2018-10-03 14:47:56 -07:00
R.B. Boyer	491826ddbc	cli: forward SIGTERM to child process of 'lock' and 'watch' subcommands (#4737 ) cli: forward SIGTERM to child process of 'lock' and 'watch' subcommands on unix This also removes the signal handler for SIGKILL as it's impossible to receive these signals.	2018-10-02 15:57:21 -05:00
Alex Dadgar	43d0f96c42	do not bootstrap with non voters	2018-09-19 17:41:36 -07:00
Kyle Havlovitz	57deb28ade	connect/ca: tighten up the intermediate signing verification	2018-09-14 16:08:54 -07:00
Kyle Havlovitz	2919519665	connect/ca: add intermediate functions to Vault ca provider	2018-09-13 13:38:32 -07:00
Kyle Havlovitz	52e8652ac5	connect/ca: add intermediate functions to Consul CA provider	2018-09-13 13:09:21 -07:00
Kyle Havlovitz	d515d25856	Merge pull request #4644 from hashicorp/ca-refactor connect/ca: rework initialization/root generation in providers	2018-09-13 13:08:34 -07:00
mkeeler	48d287ef69	Release v1.2.3	2018-09-13 15:22:25 +00:00
Paul Banks	74f2a80a42	Fix CA pruning when CA config uses string durations. (#4669 ) * Fix CA pruning when CA config uses string durations. The tl;dr here is: - Configuring LeafCertTTL with a string like "72h" is how we do it by default and should be supported - Most of our tests managed to escape this by defining them as time.Duration directly - Out actual default value is a string - Since this is stored in a map[string]interface{} config, when it is written to Raft it goes through a msgpack encode/decode cycle (even though it's written from server not over RPC). - msgpack decode leaves the string as a `[]uint8` - Some of our parsers required string and failed - So after 1 hour, a default configured server would throw an error about pruning old CAs - If a new CA was configured that set LeafCertTTL as a time.Duration, things might be OK after that, but if a new CA was just configured from config file, intialization would cause same issue but always fail still so would never prune the old CA. - Mostly this is just a janky error that got passed tests due to many levels of complicated encoding/decoding. tl;dr of the tl;dr: Yay for type safety. Map[string]interface{} combined with msgpack always goes wrong but we somehow get bitten every time in a new way :D We already fixed this once! The main CA config had the same problem so @kyhavlov already wrote the mapstructure DecodeHook that fixes it. It wasn't used in several places it needed to be and one of those is notw in `structs` which caused a dependency cycle so I've moved them. This adds a whole new test thta explicitly tests the case that broke here. It also adds tests that would have failed in other places before (Consul and Vaul provider parsing functions). I'm not sure if they would ever be affected as it is now as we've not seen things broken with them but it seems better to explicitly test that and support it to not be bitten a third time! * Typo fix * Fix bad Uint8 usage	2018-09-13 15:43:00 +01:00
Hans Hasselberg	8e235a72b4	Allow disabling the HTTP API again. (#4655 ) If you provide an invalid HTTP configuration consul will still start again instead of failing. But if you do so the build-in proxy won't be able to start which you might need for connect.	2018-09-13 16:06:04 +02:00
Kyle Havlovitz	5c7fbc284d	connect/ca: hash the consul provider ID and include isRoot	2018-09-12 13:44:15 -07:00
Pierre Souchay	1a906ef34e	Fix more unstable tests in agent and command	2018-09-12 14:49:27 +01:00
Kyle Havlovitz	c112a72880	connect/ca: some cleanup and reorganizing of the new methods	2018-09-11 16:43:04 -07:00
Pierre Souchay	2fe728c7bd	Ensure that Proxies ARE always cleaned up, event with DeregisterCriticalServiceAfter (#4649 ) This fixes https://github.com/hashicorp/consul/issues/4648	2018-09-11 17:34:09 +01:00
Matt Keeler	d3ee66eed4	Add ECS option to EDNS responses where appropriate (#4647 ) This implements parts of RFC 7871 where Consul is acting as an authoritative name server (or forwarding resolver when recursors are configured) If ECS opt is present in the request we will mirror it back and return a response with a scope of 0 (global) or with the same prefix length as the request (indicating its valid specifically for that subnet). We only mirror the prefix-length (non-global) for prepared queries as those could potentially use nearness checks that could be affected by the subnet. In the future we could get more sophisticated with determining the scope bits and allow for better caching of prepared queries that don’t rely on nearness checks. The other thing this does not do is implement the part of the ECS RFC related to originating ECS headers when acting as a intermediate DNS server (forwarding resolver). That would take a quite a bit more effort and in general provide very little value. Consul will currently forward the ECS headers between recursors and the clients transparently, we just don't originate them for non-ECS clients to get potentially more accurate "location aware" results.	2018-09-11 09:37:46 -04:00
Pierre Souchay	22500f242e	Fix unstable tests in agent, api, and command/watch	2018-09-10 16:58:53 +01:00
Mitchell Hashimoto	49b165965d	Merge pull request #4642 from hashicorp/f-ui-meta agent: aggregate service instance meta for UI purposes	2018-09-07 17:36:23 -07:00
Mitchell Hashimoto	b95348c4b1	agent: ExternalSources instead of Meta	2018-09-07 10:06:55 -07:00
Matt Keeler	cc8327ed9a	Ensure that errors setting up the DNS servers get propagated back to the shell (#4598 ) Fixes: #4578 Prior to this fix if there was an error binding to ports for the DNS servers the error would be swallowed by the gated log writer and never output. This fix propagates the DNS server errors back to the shell with a multierror.	2018-09-07 10:48:29 -04:00
Pierre Souchay	eddcf228ea	Implementation of Weights Data structures (#4468 ) * Implementation of Weights Data structures Adding this datastructure will allow us to resolve the issues #1088 and #4198 This new structure defaults to values: ``` { Passing: 1, Warning: 0 } ``` Which means, use weight of 0 for a Service in Warning State while use Weight 1 for a Healthy Service. Thus it remains compatible with previous Consul versions. * Implemented weights for DNS SRV Records * DNS properly support agents with weight support while server does not (backwards compatibility) * Use Warning value of Weights of 1 by default When using DNS interface with only_passing = false, all nodes with non-Critical healthcheck used to have a weight value of 1. While having weight.Warning = 0 as default value, this is probably a bad idea as it breaks ascending compatibility. Thus, we put a default value of 1 to be consistent with existing behaviour. * Added documentation for new weight field in service description * Better documentation about weights as suggested by @banks * Return weight = 1 for unknown Check states as suggested by @banks * Fixed typo (of -> or) in error message as requested by @mkeeler * Fixed unstable unit test TestRetryJoin * Fixed unstable tests * Fixed wrong Fatalf format in `testrpc/wait.go` * Added notes regarding DNS SRV lookup limitations regarding number of instances * Documentation fixes and clarification regarding SRV records with weights as requested by @banks * Rephrase docs	2018-09-07 15:30:47 +01:00
Kyle Havlovitz	546bdf8663	connect/ca: add Configure/GenerateRoot to provider interface	2018-09-06 19:18:59 -07:00
Mitchell Hashimoto	e9ea190df0	agent: aggregate service instance meta for UI purposes	2018-09-06 12:19:05 -07:00
Mitchell Hashimoto	99eb154f6f	agent: configure k8s go-discover	2018-09-05 13:38:13 -07:00
Martin	feb3ce4ee0	Use target service name instead of ID as connect proxy service name (#4620 )	2018-09-05 20:33:17 +01:00
Pierre Souchay	9a2ae6e8eb	Fixed more flaky tests in ./agent/consul (#4617 )	2018-09-04 14:02:47 +01:00
Pierre Souchay	92acdaa94c	Fixed flaky tests (#4626 )	2018-09-04 12:31:51 +01:00
Siva Prasad	ca35d04472	Adds a new command line flag -log-file for file based logging. (#4581 ) * Added log-file flag to capture Consul logs in a user specified file * Refactored code. * Refactored code. Added flags to rotate logs based on bytes and duration * Added the flags for log file and log rotation on the webpage * Fixed TestSantize from failing due to the addition of 3 flags * Introduced changes : mutex, data-dir log writes, rotation logic * Added test for logfile and updated the default log destination for docs * Log name now uses UnixNano * TestLogFile is now uses t.Parallel() * Removed unnecessary int64Val function * Updated docs to reflect default log name for log-file * No longer writes to data-dir and adds .log if the filename has no extension	2018-08-29 16:56:58 -04:00
Freddy	d7a404f2ee	Bugfix: Use "%#v" when formatting structs (#4600 )	2018-08-28 12:37:34 -04:00
Siva Prasad	b1a34f899f	TestAgentAntiEntropy: Wait until Consul service is up on the agent. (#4591 ) * Anti-Entropy test wait for Consul service added * Reverted some tests back to using WaitForLeader	2018-08-28 09:52:11 -04:00
Pierre Souchay	5e0218ccf4	Fix unit test TestOperatorAutopilotGetConfigCommand (#4594 )	2018-08-27 13:29:25 -04:00
Pierre Souchay	aea31d3c5d	Fixed unstable test TestUiNodeInfo (#4586 )	2018-08-27 11:49:14 -04:00
Pierre Souchay	b898131723	[BUGFIX] Avoid returning empty data on startup of a non-leader server (#4554 ) Ensure that DB is properly initialized when performing stale queries Addresses: - https://github.com/hashicorp/consul-replicate/issues/82 - https://github.com/hashicorp/consul/issues/3975 - https://github.com/hashicorp/consul-template/issues/1131	2018-08-23 12:06:39 -04:00
Miroslav Bagljas	3c23979afd	Fixes #4483 : Add support for Authorization: Bearer token Header (#4502 ) Added Authorization Bearer token support as per RFC6750 * appended Authorization header token parsing after X-Consul-Token * added test cases * updated website documentation to mention Authorization header * improve tests, improve Bearer parsing	2018-08-17 16:18:42 -04:00
Matt Keeler	e81c85c051	Fix #4515 : Segfault when serf_wan port was -1 but reconnect_time_wan was set (#4531 ) Fixes #4515 This just slightly refactors the logic to only attempt to set the serf wan reconnect timeout when the rest of the serf wan settings are configured - thus avoiding a segfault.	2018-08-17 14:44:25 -04:00
Kyle Havlovitz	e5e1f867e5	Merge branch 'master' into ca-snapshot-fix	2018-08-16 13:00:54 -07:00
Kyle Havlovitz	f186edc42c	fsm: add connect service config to snapshot/restore test	2018-08-16 12:58:54 -07:00
nickmy9729	beddf03b26	Added code to allow snapshot inclusion of NodeMeta (#4527 )	2018-08-16 15:33:35 -04:00
Kyle Havlovitz	b51d76f469	fsm: add missing CA config to snapshot/restore logic	2018-08-16 11:58:50 -07:00
Kyle Havlovitz	4b35d877ca	autopilot: don't follow the normal server removal rules for nonvoters	2018-08-14 14:24:51 -07:00
Kyle Havlovitz	ea14482376	Fix stats fetcher healthcheck RPCs not being independent	2018-08-14 14:23:52 -07:00
Pierre Souchay	0d6de257a2	Display more information about check being not properly added when it fails (#4405 ) * Display more information about check being not properly added when it fails It follows an incident where we add lots of error messages: [WARN] consul.fsm: EnsureRegistration failed: failed inserting check: Missing service registration That seems related to Consul failing to restart on respective agents. Having Node information as well as service information would help diagnose the issue. * Renamed ensureCheckIfNodeMatches() as requested by @banks	2018-08-14 17:45:33 +01:00
Freddy	6d43d24edb	Improve reliability of tests with TestAgent (#4525 ) - Add WaitForTestAgent to tests flaky due to missing serfHealth registration - Fix bug in retries calling Fatalf with *testing.T - Convert TestLockCommand_ChildExitCode to table driven test	2018-08-14 12:08:33 -04:00
Pierre Souchay	ef3b81ab13	Allow to rename nodes with IDs, will fix #3974 and #4413 (#4415 ) * Allow to rename nodes with IDs, will fix #3974 and #4413 This change allow to rename any well behaving recent agent with an ID to be renamed safely, ie: without taking the name of another one with case insensitive comparison. Deprecated behaviour warning ---------------------------- Due to asceding compatibility, it is still possible however to "take" the name of another name by not providing any ID. Note that when not providing any ID, it is possible to have 2 nodes having similar names with case differences, ie: myNode and mynode which might lead to DB corruption on Consul server side and lead to server not properly restarting. See #3983 and #4399 for Context about this change. Disabling registration of nodes without IDs as specified in #4414 should probably be the way to go eventually. * Removed the case-insensitive search when adding a node within the else block since it breaks the test TestAgentAntiEntropy_Services While the else case is probably legit, it will be fixed with #4414 in a later release. * Added again the test in the else to avoid duplicated names, but enforce this test only for nodes having IDs. Thus most tests without any ID will work, and allows us fixing * Added more tests regarding request with/without IDs. `TestStateStore_EnsureNode` now test registration and renaming with IDs `TestStateStore_EnsureNodeDeprecated` tests registration without IDs and tests removing an ID from a node as well as updated a node without its ID (deprecated behaviour kept for backwards compatibility) * Do not allow renaming in case of conflict, including when other node has no ID * Fixed function GetNodeID that was not working due to wrong type when searching node from its ID Thus, all tests about renaming were not working properly. Added the full test cas that allowed me to detect it. * Better error messages, more tests when nodeID is not a valid UUID in GetNodeID() * Added separate TestStateStore_GetNodeID to test GetNodeID. More complete test coverage for GetNodeID * Added new unit test `TestStateStore_ensureNoNodeWithSimilarNameTxn` Also fixed comments to be clearer after remarks from @banks * Fixed error message in unit test to match test case * Use uuid.ParseUUID to parse Node.ID as requested by @mkeeler	2018-08-10 11:30:45 -04:00
Siva Prasad	c88900aaa9	PR to fix TestAgent_IndexChurn and TestPreparedQuery_Wrapper. (#4512 ) * Fixes TestAgent_IndexChurn * Fixes TestPreparedQuery_Wrapper * Increased sleep in agent_test for IndexChurn to 500ms * Made the comment about joinWAN operation much less of a cliffhanger	2018-08-09 12:40:07 -04:00
Armon Dadgar	4f1fd34e9e	consul: Update buffer sizes	2018-08-08 10:26:58 -07:00
Siva Prasad	288d350a73	Revert "CA initialization while boostrapping and TestLeader_ChangeServerID fix." (#4497 ) * Revert "BUGFIX: Unit test relying on WaitForLeader() did not work due to wrong test (#4472)" This reverts commit `cec5d72396`. * Revert "CA initialization while boostrapping and TestLeader_ChangeServerID fix. (#4493)" This reverts commit `589b589b53`.	2018-08-07 08:29:48 -04:00
Pierre Souchay	cec5d72396	BUGFIX: Unit test relying on WaitForLeader() did not work due to wrong test (#4472 ) - Improve resilience of testrpc.WaitForLeader() - Add additionall retry to CI - Increase "go test" timeout to 8m - Add wait for cluster leader to several tests in the agent package - Add retry to some tests in the api and command packages	2018-08-06 19:46:09 -04:00

1 2 3 4 5 ...

1226 Commits