futriix

Author	SHA1	Message	Date
Oran Agra	573246f73c	if diskless repl child is killed, make sure to reap the pid (#7742 ) Starting redis 6.0 and the changes we made to the diskless master to be suitable for TLS, I made the master avoid reaping (wait3) the pid of the child until we know all replicas are done reading their rdb. I did that in order to avoid a state where the rdb_child_pid is -1 but we don't yet want to start another fork (still busy serving that data to replicas). It turns out that the solution used so far was problematic in case the fork child was being killed (e.g. by the kernel OOM killer), in that case there's a chance that we currently disabled the read event on the rdb pipe, since we're waiting for a replica to become writable again. and in that scenario the master would have never realized the child exited, and the replica will remain hung too. Note that there's no mechanism to detect a hung replica while it's in rdb transfer state. The solution here is to add another pipe which is used by the parent to tell the child it is safe to exit. this mean that when the child exits, for whatever reason, it is safe to reap it. Besides that, i'm re-introducing an adjustment to REPLCONF ACK which was part of #6271 (Accelerate diskless master connections) but was dropped when that PR was rebased after the TLS fork/pipe changes (5a47794). Now that RdbPipeCleanup no longer calls checkChildrenDone, and the ACK has chance to detect that the child exited, it should be the one to call it so that we don't have to wait for cron (server.hz) to do that.	2020-09-06 16:43:57 +03:00
Oran Agra	9ef8d2f671	Run active defrag while blocked / loading (#7726 ) During long running scripts or loading RDB/AOF, we may need to do some defragging. Since processEventsWhileBlocked is called periodically at unknown intervals, and many cron jobs either depend on run_with_period (including active defrag), or rely on being called at server.hz rate (i.e. active defrag knows ho much time to run by looking at server.hz), the whileBlockedCron may have to run a loop triggering the cron jobs in it (currently only active defrag) several times. Other changes: - Adding a test for defrag during aof loading. - Changing key-load-delay config to take negative values for fractions of a microsecond sleep	2020-09-03 08:47:29 +03:00
Yossi Gottlieb	f6d04d01b9	Add oom-score-adj configuration option to control Linux OOM killer. (#1690 ) Add Linux kernel OOM killer control option. This adds the ability to control the Linux OOM killer oom_score_adj parameter for all Redis processes, depending on the process role (i.e. master, replica, background child). A oom-score-adj global boolean flag control this feature. In addition, specific values can be configured using oom-score-adj-values if additional tuning is required. (cherry picked from commit 2530dc0ebd8be8d792f4673073401377cd5bdc42)	2020-09-01 09:27:58 +03:00
Arun Ranganathan	e81bac32fd	Show threading configuration in INFO output (#7446 ) Co-authored-by: Oran Agra <oran@redislabs.com> (cherry picked from commit f6cad30bb69b2ad35bb0a870077fac2d4605d727)	2020-09-01 09:27:58 +03:00
Meir Shpilraien (Spielrein)	2257f38b68	This PR introduces a new loaded keyspace event (#7536 ) Co-authored-by: Oran Agra <oran@redislabs.com> Co-authored-by: Itamar Haber <itamar@redislabs.com> (cherry picked from commit 8d826393191399e132bd9e56fb51ed83223cc5ca)	2020-09-01 09:27:58 +03:00
Oran Agra	2165a78d10	Fix rejectCommand trims newline in shared error objects, hung clients (#7714 ) 65a3307bc (released in 6.0.6) has a side effect, when processCommand rejects a command with pre-made shared object error string, it trims the newlines from the end of the string. if that string is later used with addReply, the newline will be missing, breaking the protocol, and leaving the client hung. It seems that the only scenario which this happens is when replying with -LOADING to some command, and later using that reply from the CONFIG SET command (still during loading). this will result in hung client. Refactoring the code in order to avoid trimming these newlines from shared string objects, and do the newline trimming only in other cases where it's needed. Co-authored-by: Guy Benoish <guy.benoish@redislabs.com> (cherry picked from commit 9fcd9e191e6f54276688fb7c74e1d5c3c4be9a75)	2020-09-01 09:27:58 +03:00
杨博东	113d5ae872	Fix flock cluster config may cause failure to restart after kill -9 (#7674 ) After fork, the child process(redis-aof-rewrite) will get the fd opened by the parent process(redis), when redis killed by kill -9, it will not graceful exit(call prepareForShutdown()), so redis-aof-rewrite thread may still alive, the fd(lock) will still be held by redis-aof-rewrite thread, and redis restart will fail to get lock, means fail to start. This issue was causing failures in the cluster tests in github actions. Co-authored-by: Oran Agra <oran@redislabs.com> (cherry picked from commit cbaf3c5bbafd43e009a2d6b38dd0e9fc450a3e12)	2020-09-01 09:27:58 +03:00
Jiayuan Chen	096285ab64	Add optional tls verification (#7502 ) Adds an `optional` value to the previously boolean `tls-auth-clients` configuration keyword. Co-authored-by: Yossi Gottlieb <yossigo@gmail.com> (cherry picked from commit f31260b0445f5649449da41555e1272a40ae4af7)	2020-09-01 09:27:58 +03:00
grishaf	395d339234	Fix prepareForShutdown function declaration (#7566 ) (cherry picked from commit 4126ca466fb9fbf0d84a66e31c9f59f28e2fbff6)	2020-09-01 09:27:58 +03:00
Oran Agra	9fcd9e191e	Fix rejectCommand trims newline in shared error objects, hung clients (#7714 ) 65a3307bc (released in 6.0.6) has a side effect, when processCommand rejects a command with pre-made shared object error string, it trims the newlines from the end of the string. if that string is later used with addReply, the newline will be missing, breaking the protocol, and leaving the client hung. It seems that the only scenario which this happens is when replying with -LOADING to some command, and later using that reply from the CONFIG SET command (still during loading). this will result in hung client. Refactoring the code in order to avoid trimming these newlines from shared string objects, and do the newline trimming only in other cases where it's needed. Co-authored-by: Guy Benoish <guy.benoish@redislabs.com>	2020-08-27 12:54:01 +03:00
Oran Agra	8bdcbbb085	Update memory metrics for INFO during loading (#7690 ) During a long AOF or RDB loading, the memory stats were not updated, and INFO would return stale data, specifically about fragmentation and RSS. In the past some of these were sampled directly inside the INFO command, but were moved to cron as an optimization. This commit introduces a concept of loadingCron which should take some of the responsibilities of serverCron. It attempts to limit it's rate to approximately the server Hz, but may not be very accurate. In order to avoid too many system call, we use the cached ustime, and also make sure to update it in both AOF loading and RDB loading inside processEventsWhileBlocked (it seems AOF loading was missing it).	2020-08-27 11:09:32 +03:00
John Sully	d1f45b8ebf	Implement use-fork config (fails with diskless repl) Former-commit-id: f2d5c2bca22e9fd506db123c47b7f60cdded7e2c	2020-08-24 03:17:59 +00:00
杨博东	cbaf3c5bba	Fix flock cluster config may cause failure to restart after kill -9 (#7674 ) After fork, the child process(redis-aof-rewrite) will get the fd opened by the parent process(redis), when redis killed by kill -9, it will not graceful exit(call prepareForShutdown()), so redis-aof-rewrite thread may still alive, the fd(lock) will still be held by redis-aof-rewrite thread, and redis restart will fail to get lock, means fail to start. This issue was causing failures in the cluster tests in github actions. Co-authored-by: Oran Agra <oran@redislabs.com>	2020-08-20 08:59:02 +03:00
John Sully	035cde502a	Fast cleanup of snapshots without leaving them forever Former-commit-id: fdd83b2b49244ed2988b080892ee5cffe9fd2684	2020-08-17 00:33:37 +00:00
John Sully	b71f16471b	Allow garbage collection of generic data Former-commit-id: feadb7fb1845027422bcfca43dbcb6097409b8dc	2020-08-17 00:32:48 +00:00
John Sully	6ba6652fcf	Don't try and consolidate snapshots with a depth of 1 Former-commit-id: 26c298bd9bc4e2c6981de5c20284120ea54580c3	2020-08-16 00:26:05 +00:00
John Sully	71506d7e0a	Don't free snapshot objects in a critical path (under the AE lock) Former-commit-id: d0da3d3cb74334cc8a2d14f4bdaef7935181700a	2020-08-16 00:13:19 +00:00
John Sully	34937b0ad5	Rehash efficiency Former-commit-id: fab383156626ec683881101c22eb2f6c2cea4c5d	2020-08-15 23:05:56 +00:00
John Sully	07c019fd3d	Prevent unnecessary copy when overwriting a value from a snapshot Former-commit-id: 654a7bc6ea82f4ac45a1c1a25c794e1c27c0d902	2020-08-15 22:59:01 +00:00
John Sully	db193a1ef1	Merge branch 'unstable' into keydbpro Former-commit-id: ae482585f0dc470efd73833f74111c2f87a172c5	2020-08-15 21:29:00 +00:00
Yossi Gottlieb	2530dc0ebd	Add oom-score-adj configuration option to control Linux OOM killer. (#1690 ) Add Linux kernel OOM killer control option. This adds the ability to control the Linux OOM killer oom_score_adj parameter for all Redis processes, depending on the process role (i.e. master, replica, background child). A oom-score-adj global boolean flag control this feature. In addition, specific values can be configured using oom-score-adj-values if additional tuning is required.	2020-08-12 17:58:56 +03:00
Tyson Andre	6f11acbd67	Implement SMISMEMBER key member [member ...] (#7615 ) This is a rebased version of #3078 originally by shaharmor with the following patches by TysonAndre made after rebasing to work with the updated C API: 1. Add 2 more unit tests (wrong argument count error message, integer over 64 bits) 2. Use addReplyArrayLen instead of addReplyMultiBulkLen. 3. Undo changes to src/help.h - for the ZMSCORE PR, I heard those should instead be automatically generated from the redis-doc repo if it gets updated Motivations: - Example use case: Client code to efficiently check if each element of a set of 1000 items is a member of a set of 10 million items. (Similar to reasons for working on #7593) - HMGET and ZMSCORE already exist. This may lead to developers deciding to implement functionality that's best suited to a regular set with a data type of sorted set or hash map instead, for the multi-get support. Currently, multi commands or lua scripting to call sismember multiple times would almost definitely be less efficient than a native smismember for the following reasons: - Need to fetch the set from the string every time instead of reusing the C pointer. - Using pipelining or multi-commands would result in more bytes sent and received by the client for the repeated SISMEMBER KEY sections. - Need to specially encode the data and decode it from the client for lua-based solutions. - Proposed solutions using Lua or SADD/SDIFF could trigger writes to memory, which is undesirable on a redis replica server or when commands get replicated to replicas. Co-Authored-By: Shahar Mor <shahar@peer5.com> Co-Authored-By: Tyson Andre <tysonandre775@hotmail.com>	2020-08-11 11:55:06 +03:00
Rajat Pawar	59d437c727	Fix comment about ACLGetCommandPerm()	2020-08-10 23:11:26 -07:00
杨博东	229327ad8b	Avoid redundant calls to signalKeyAsReady (#7625 ) signalKeyAsReady has some overhead (namely dictFind) so we should only call it when there are clients blocked on the relevant type (BLOCKED_*)	2020-08-11 08:18:09 +03:00
John Sully	e87dee8dc7	Add build flag to disable MVCC tstamps Former-commit-id: f17d178d03f44abcdaddd851a313dd3f7ec87ed5	2020-08-10 06:10:24 +00:00
John Sully	649745924b	RocksDB Read Performance Improvements Former-commit-id: 80cca4869888e048e10e11f1f20796c482c3e5b3	2020-08-09 23:36:20 +00:00
Wang Yuan	1ef014ee6b	Fix applying zero offset to null pointer when creating moduleFreeContextReusedClient (#7323 ) Before this fix we where attempting to select a db before creating db the DB, see: #7323 This issue doesn't seem to have any implications, since the selected DB index is 0, the db pointer remains NULL, and will later be correctly set before using this dummy client for the first time. As we know, we call 'moduleInitModulesSystem()' before 'initServer()'. We will allocate memory for server.db in 'initServer', but we call 'createClient()' that will call 'selectDb()' in 'moduleInitModulesSystem()', before the databases where created. Instead, we should call 'createClient()' for moduleFreeContextReusedClient after 'initServer()'.	2020-08-08 14:36:41 +03:00
Oran Agra	c17e597d05	Accelerate diskless master connections, and general re-connections (#6271 ) Diskless master has some inherent latencies. 1) fork starts with delay from cron rather than immediately 2) replica is put online only after an ACK. but the ACK was sent only once a second. 3) but even if it would arrive immediately, it will not register in case cron didn't yet detect that the fork is done. Besides that, when a replica disconnects, it doesn't immediately attempts to re-connect, it waits for replication cron (one per second). in case it was already online, it may be important to try to re-connect as soon as possible, so that the backlog at the master doesn't vanish. In case it disconnected during rdb transfer, one can argue that it's not very important to re-connect immediately, but this is needed for the "diskless loading short read" test to be able to run 100 iterations in 5 seconds, rather than 3 (waiting for replication cron re-connection) changes in this commit: 1) sync command starts a fork immediately if no sync_delay is configured 2) replica sends REPLCONF ACK when done reading the rdb (rather than on 1s cron) 3) when a replica unexpectedly disconnets, it immediately tries to re-connect rather than waiting 1s 4) when when a child exits, if there is another replica waiting, we spawn a new one right away, instead of waiting for 1s replicationCron. 5) added a call to connectWithMaster from replicationSetMaster. which is called from the REPLICAOF command but also in 3 places in cluster.c, in all of these the connection attempt will now be immediate instead of delayed by 1 second. side note: we can add a call to rdbPipeReadHandler in replconfCommand when getting a REPLCONF ACK from the replica to solve a race where the replica got the entire rdb and EOF marker before we detected that the pipe was closed. in the test i did see this race happens in one about of some 300 runs, but i concluded that this race is unlikely in real life (where the replica is on another host and we're more likely to first detect the pipe was closed. the test runs 100 iterations in 3 seconds, so in some cases it'll take 4 seconds instead (waiting for another REPLCONF ACK). Removing unneeded startBgsaveForReplication from updateSlavesWaitingForBgsave Now that CheckChildrenDone is calling the new replicationStartPendingFork (extracted from serverCron) there's actually no need to call startBgsaveForReplication from updateSlavesWaitingForBgsave anymore, since as soon as updateSlavesWaitingForBgsave returns, CheckChildrenDone is calling replicationStartPendingFork that handles that anyway. The code in updateSlavesWaitingForBgsave had a bug in which it ignored repl-diskless-sync-delay, but removing that code shows that this bug was hiding another bug, which is that the max_idle should have used >= and not >, this one second delay has a big impact on my new test.	2020-08-06 16:53:06 +03:00
Oran Agra	90b717e723	Assertion and panic, print crash log without generating SIGSEGV This makes it possible to add tests that generate assertions, and run them with valgrind, making sure that there are no memory violations prior to the assertion. New config options: - crash-log-enabled - can be disabled for cleaner core dumps - crash-memcheck-enabled - useful for faster termination after a crash - use-exit-on-panic - to be used by the test suite so that valgrind can detect leaks and memory corruptions Other changes: - Crash log is printed even on system that dont HAVE_BACKTRACE, i.e. in both SIGSEGV and assert / panic - Assertion and panic won't print registers and code around EIP (which was useless), but will do fast memory test (which may still indicate that the assertion was due to memory corrpution) I had to reshuffle code in order to re-use it, so i extracted come code into function without actually doing any changes to the code: - logServerInfo - logModulesInfo - doFastMemoryTest (with the exception of it being conditional) - dumpCodeAroundEIP changes to the crash report on segfault: - logRegisters is called right after the stack trace (before info) done just in order to have more re-usable code - stack trace skips the first two items on the stack (the crash log and signal handler functions)	2020-08-06 16:47:27 +03:00
Tyson Andre	f11f26cc53	Add a ZMSCORE command returning an array of scores. (#7593 ) Syntax: `ZMSCORE KEY MEMBER [MEMBER ...]` This is an extension of #2359 amended by Tyson Andre to work with the changed unstable API, add more tests, and consistently return an array. - It seemed as if it would be more likely to get reviewed after updating the implementation. Currently, multi commands or lua scripting to call zscore multiple times would almost definitely be less efficient than a native ZMSCORE for the following reasons: - Need to fetch the set from the string every time instead of reusing the C pointer. - Using pipelining or multi-commands would result in more bytes sent by the client for the repeated `ZMSCORE KEY` sections. - Need to specially encode the data and decode it from the client for lua-based solutions. - The fastest solution I've seen for large sets(thousands or millions) involves lua and a variadic ZADD, then a ZINTERSECT, then a ZRANGE 0 -1, then UNLINK of a temporary set (or lua). This is still inefficient. Co-authored-by: Tyson Andre <tysonandre775@hotmail.com>	2020-08-04 17:49:33 +03:00
Arun Ranganathan	f6cad30bb6	Show threading configuration in INFO output (#7446 ) Co-authored-by: Oran Agra <oran@redislabs.com>	2020-07-29 08:46:44 +03:00
Jiayuan Chen	f31260b044	Add optional tls verification (#7502 ) Adds an `optional` value to the previously boolean `tls-auth-clients` configuration keyword. Co-authored-by: Yossi Gottlieb <yossigo@gmail.com>	2020-07-28 10:45:21 +03:00
grishaf	4126ca466f	Fix prepareForShutdown function declaration (#7566 )	2020-07-26 08:27:30 +03:00
Meir Shpilraien (Spielrein)	8d82639319	This PR introduces a new loaded keyspace event (#7536 ) Co-authored-by: Oran Agra <oran@redislabs.com> Co-authored-by: Itamar Haber <itamar@redislabs.com>	2020-07-23 12:38:51 +03:00
Yossi Gottlieb	7a536c2912	TLS: Session caching configuration support. (#7420 ) * TLS: Session caching configuration support. * TLS: Remove redundant config initialization. (cherry picked from commit 3e6f2b1a45176ac3d81b95cb6025f30d7aaa1393)	2020-07-20 21:08:26 +03:00
Oran Agra	95ba01b538	RESTORE ABSTTL won't store expired keys into the db (#7472 ) Similarly to EXPIREAT with TTL in the past, which implicitly deletes the key and return success, RESTORE should not store key that are already expired into the db. When used together with REPLACE it should emit a DEL to keyspace notification and replication stream. (cherry picked from commit 5977a94842a25140520297fe4bfda15e0e4de711)	2020-07-20 21:08:26 +03:00
Oran Agra	05e483cbb3	EXEC always fails with EXECABORT and multi-state is cleared In order to support the use of multi-exec in pipeline, it is important that MULTI and EXEC are never rejected and it is easy for the client to know if the connection is still in multi state. It was easy to make sure MULTI and DISCARD never fail (done by previous commits) since these only change the client state and don't do any actual change in the server, but EXEC is a different story. Since in the past, it was possible for clients to handle some EXEC errors and retry the EXEC, we now can't affort to return any error on EXEC other than EXECABORT, which now carries with it the real reason for the abort too. Other fixes in this commit: - Some checks that where performed at the time of queuing need to be re- validated when EXEC runs, for instance if the transaction contains writes commands, it needs to be aborted. there was one check that was already done in execCommand (-READONLY), but other checks where missing: -OOM, -MISCONF, -NOREPLICAS, -MASTERDOWN - When a command is rejected by processCommand it was rejected with addReply, which was not recognized as an error in case the bad command came from the master. this will enable to count or MONITOR these errors in the future. - make it easier for tests to create additional (non deferred) clients. - add tests for the fixes of this commit. (cherry picked from commit 65a3307bc95aadbc91d85cdf9dfbe1b3493222ca)	2020-07-20 21:08:26 +03:00
John Sully	655bc912e6	Add new storage-provider-options config Former-commit-id: 195a28beecc6094f959ddafef7fe33f5b55e4047	2020-07-14 04:24:46 +00:00
John Sully	782e675072	Remove unnecessary work from critical path (Latency Fixes) Former-commit-id: 096a90deb7afe489875d3186f3f8f43e41fea329	2020-07-13 16:03:52 +00:00
John Sully	91a803815d	Merge branch 'flash_cache' into keydbpro Former-commit-id: 2a721ef645921d62b39f1374c0a3f5c92b00fae5	2020-07-13 03:53:34 +00:00
John Sully	6a57593467	Merge branch 'PRO_RELEASE_6' into keydbpro Former-commit-id: bffe010ea5279bee869bc61cc6d933979e10bbea	2020-07-13 03:32:14 +00:00
John Sully	1ad2d96697	Merge branch 'unstable' into keydbpro Former-commit-id: 0dafbc254a0efd5ee302d5c58fb2ca0a85110104	2020-07-13 03:31:47 +00:00
John Sully	cfcb5ac5c7	Add the KEYDB.MEXISTS command, see issue #203 Former-commit-id: 5619f515285b08d9c443425de1f3092ae3058d40	2020-07-12 21:42:11 +00:00
John Sully	c5f6cb1ba5	Add multi-master-no-forward command to reduce bus traffic with multi-master Former-commit-id: d99d06b1250a51ea4bc54f678f451acbb7901e33	2020-07-12 19:25:19 +00:00
John Sully	17661f2382	Implement storage key cache, and writeback memory model Former-commit-id: 732bd9c153459f1174475ad67de36c399ddbe359	2020-07-11 21:23:48 +00:00
Yossi Gottlieb	3e6f2b1a45	TLS: Session caching configuration support. (#7420 ) * TLS: Session caching configuration support. * TLS: Remove redundant config initialization.	2020-07-10 11:33:47 +03:00
Oran Agra	5977a94842	RESTORE ABSTTL won't store expired keys into the db (#7472 ) Similarly to EXPIREAT with TTL in the past, which implicitly deletes the key and return success, RESTORE should not store key that are already expired into the db. When used together with REPLACE it should emit a DEL to keyspace notification and replication stream.	2020-07-10 10:02:37 +03:00
John Sully	52123f9529	Merge branch 'keydbpro' into PRO_RELEASE_6 Former-commit-id: 243dcb3853cc965109cb24a940229db7844cdd11	2020-07-10 04:11:57 +00:00
John Sully	f03b75b005	MVCC scan support filtering by type on the async thread Former-commit-id: 14f8c0ff686b93976eead5fa6bf526c2eecb5ae0	2020-07-10 03:43:56 +00:00
John Sully	34482220af	Fix issue where SCAN misses elements while snapshot is in flight Former-commit-id: ce005d748ebf0e116d674a96f74d698d17394010	2020-07-10 01:43:51 +00:00

... 3 4 5 6 7 ...

1209 Commits