futriix

Author	SHA1	Message	Date
Eran Liberty	b120366d48	Allow exec with read commands on readonly replica in cluster (#7766 ) There was a bug. Although cluster replicas would allow read commands, they would not allow a MULTI-EXEC that's composed solely of read commands. Adds tests for coverage. Co-authored-by: Oran Agra <oran@redislabs.com> Co-authored-by: Eran Liberty <eranl@amazon.com>	2020-09-09 09:35:42 +03:00
Oran Agra	0a1e734193	handle cur_test for nested tests if there are nested tests and nested servers, we need to restore the previous value of cur_test when a test exist. example: ``` test{test 1} { start_server { test{test 1.1 - master only} { } start_server { test{test 1.2 - with replication} { } } } } ``` when `test 1.1 - master only exists`, we're still inside `test 1`	2020-09-08 14:12:03 +03:00
bodong.ybd	f22fa9594d	Tests: Some fixes for macOS 1) cur_test: when restart_server, "no such variable" error occurs ./runtest --single integration/rdb test {client freed during loading} SET ::cur_test restart_server kill_server test "Check for memory leaks (pid $pid)" SET ::cur_test UNSET ::cur_test UNSET ::cur_test // This global variable has been unset. 2) `ps --ppid` not available on macOS platform, can be replaced with `pgrep -P pid`.	2020-09-08 14:27:53 +08:00
Oran Agra	b491d477c3	Fix cluster consistency-check test (#7754 ) This test was failing from time to time see discussion at the bottom of #7635 This was probably due to timing, the DEBUG SLEEP executed by redis-cli didn't sleep for enough time. This commit changes: 1) use SET-ACTIVE-EXPIRE instead of DEBUG SLEEP 2) reduce many `after` sleeps with retry loops to speed up the test. 3) add many comment explaining the different steps of the test and it's purpose. 4) config appendonly before populating the volatile keys, so that they'll be part of the AOF command stream rather than the preamble RDB portion. other complications: recently kill_instance switched from SIGKILL to SIGTERM, and this would sometimes fail since there was an AOFRW running in the background. now we wait for it to end before attempting the kill.	2020-09-07 18:06:25 +03:00
Yossi Gottlieb	2df4cb93ac	Tests: fix unmonitored servers. (#7756 ) There is an inherent race condition in port allocation for spawned servers. If a server fails to start because a port is taken, a new port is allocated. This fixes a problem where the logs are not truncated and as a result a large number of unmonitored servers are started.	2020-09-07 17:30:36 +03:00
Oran Agra	42ba7a1b75	fix broken cluster/sentinel tests by recent commit (#7752 ) 2b998de46 added a file for stderr to keep valgrind log but i forgot to add a similar thing when valgrind isn't being used. the result is that `glob */err.txt` fails.	2020-09-07 16:26:11 +03:00
Oran Agra	573246f73c	if diskless repl child is killed, make sure to reap the pid (#7742 ) Starting redis 6.0 and the changes we made to the diskless master to be suitable for TLS, I made the master avoid reaping (wait3) the pid of the child until we know all replicas are done reading their rdb. I did that in order to avoid a state where the rdb_child_pid is -1 but we don't yet want to start another fork (still busy serving that data to replicas). It turns out that the solution used so far was problematic in case the fork child was being killed (e.g. by the kernel OOM killer), in that case there's a chance that we currently disabled the read event on the rdb pipe, since we're waiting for a replica to become writable again. and in that scenario the master would have never realized the child exited, and the replica will remain hung too. Note that there's no mechanism to detect a hung replica while it's in rdb transfer state. The solution here is to add another pipe which is used by the parent to tell the child it is safe to exit. this mean that when the child exits, for whatever reason, it is safe to reap it. Besides that, i'm re-introducing an adjustment to REPLCONF ACK which was part of #6271 (Accelerate diskless master connections) but was dropped when that PR was rebased after the TLS fork/pipe changes (5a47794). Now that RdbPipeCleanup no longer calls checkChildrenDone, and the ACK has chance to detect that the child exited, it should be the one to call it so that we don't have to wait for cron (server.hz) to do that.	2020-09-06 16:43:57 +03:00
Oran Agra	2b998de460	Improve valgrind support for cluster tests (#7725 ) - redirect valgrind reports to a dedicated file rather than console - try to avoid killing instances with SIGKILL so that we get the memory leak report (killing with SIGTERM before resorting to SIGKILL) - search for valgrind reports when done, print them and fail the tests - add --dont-clean option to keep the logs on exit - fix exit error code when crash is found (would have exited with 0) changes that affect the normal redis test suite: - refactor check_valgrind_errors into two functions one to search and one to report - move the search half into util.tcl to serve the cluster tests too - ignore "address range perms" valgrind warnings which seem non relevant.	2020-09-06 11:11:49 +03:00
Oran Agra	fe5da2e60d	test infra - add durable mode to work around test suite crashing in some cases a command that returns an error possibly due to a timing issue causes the tcl code to crash and thus prevents the rest of the tests from running. this adds an option to make the test proceed despite the crash. maybe it should be the default mode some day.	2020-09-06 09:59:19 +03:00
Oran Agra	1b7ba44e79	test infra - wait_done_loading reduce code duplication in aof.tcl. move creation of clients into the test so that it can be skipped	2020-09-06 09:59:19 +03:00
Oran Agra	b65e5aca86	test infra - flushall between tests in external mode	2020-09-06 09:59:19 +03:00
Oran Agra	677d14c213	test infra - improve test skipping ability - skip full units - skip a single test (not just a list of tests) - when skipping tag, skip spinning up servers, not just the tests - skip tags when running against an external server too - allow using multiple tags (split them)	2020-09-06 09:59:19 +03:00
Oran Agra	e3e69c25fd	test infra - reduce disk space usage this is important when running a test with --loop	2020-09-06 09:59:19 +03:00
Oran Agra	9d527d076b	test infra - write test name to logfile	2020-09-06 09:59:19 +03:00
Oran Agra	9ef8d2f671	Run active defrag while blocked / loading (#7726 ) During long running scripts or loading RDB/AOF, we may need to do some defragging. Since processEventsWhileBlocked is called periodically at unknown intervals, and many cron jobs either depend on run_with_period (including active defrag), or rely on being called at server.hz rate (i.e. active defrag knows ho much time to run by looking at server.hz), the whileBlockedCron may have to run a loop triggering the cron jobs in it (currently only active defrag) several times. Other changes: - Adding a test for defrag during aof loading. - Changing key-load-delay config to take negative values for fractions of a microsecond sleep	2020-09-03 08:47:29 +03:00
Oran Agra	4bb40a9688	Reduce the probability of failure when start redis in runtest-cluster #7554 (#7635 ) When runtest-cluster, at first, we need to create a cluster use spawn_instance, a port which is not used is choosen, however sometimes we can't run server on the port. possibley due to a race with another process taking it first. such as redis/redis/runs/896537490. It may be due to the machine problem or In order to reduce the probability of failure when start redis in runtest-cluster, we attemp to use another port when find server do not start up. Co-authored-by: Oran Agra <oran@redislabs.com> Co-authored-by: yanhui13 <yanhui13@meituan.com> (cherry picked from commit e2d64485b8262971776fb1be803c7296c98d1572)	2020-09-01 09:27:58 +03:00
Yossi Gottlieb	f38e2802b6	Fix oom-score-adj on older distros. (#7724 ) Don't assume `ps` handles `-h` to display output without headers and manually trim headers line from output. (cherry picked from commit b61b663895f16d9f559a14c408c225062254a57b)	2020-09-01 09:27:58 +03:00
Yossi Gottlieb	f6d04d01b9	Add oom-score-adj configuration option to control Linux OOM killer. (#1690 ) Add Linux kernel OOM killer control option. This adds the ability to control the Linux OOM killer oom_score_adj parameter for all Redis processes, depending on the process role (i.e. master, replica, background child). A oom-score-adj global boolean flag control this feature. In addition, specific values can be configured using oom-score-adj-values if additional tuning is required. (cherry picked from commit 2530dc0ebd8be8d792f4673073401377cd5bdc42)	2020-09-01 09:27:58 +03:00
Meir Shpilraien (Spielrein)	2fc915f509	see #7544 , added RedisModule_HoldString api. (#7577 ) Added RedisModule_HoldString that either returns a shallow copy of the given String (by increasing the String ref count) or a new deep copy of String in case its not possible to get a shallow copy. Co-authored-by: Itamar Haber <itamar@redislabs.com> (cherry picked from commit 3f494cc49d25929f27fa75a78d9921a9dee771f2)	2020-09-01 09:27:58 +03:00
Meir Shpilraien (Spielrein)	2257f38b68	This PR introduces a new loaded keyspace event (#7536 ) Co-authored-by: Oran Agra <oran@redislabs.com> Co-authored-by: Itamar Haber <itamar@redislabs.com> (cherry picked from commit 8d826393191399e132bd9e56fb51ed83223cc5ca)	2020-09-01 09:27:58 +03:00
valentinogeron	e6f6731c66	EXEC with only read commands should not be rejected when OOM (#7696 ) If the server gets MULTI command followed by only read commands, and right before it gets the EXEC it reaches OOM, the client will get OOM response. So, from now on, it will get OOM response only if there was at least one command that was tagged with `use-memory` flag (cherry picked from commit b7289e912cbe1a011a5569cd67929e83731b9660)	2020-09-01 09:27:58 +03:00
Valentino Geron	19ef1f371d	Fix LPOS command when RANK is greater than matches When calling to LPOS command when RANK is higher than matches, the return value is non valid response. For example: ``` LPUSH l a :1 LPOS l b RANK 5 COUNT 10 -4 ``` It may break client-side parser. Now, we count how many replies were replied in the array. ``` LPUSH l a :1 LPOS l b RANK 5 COUNT 10 0 ``` (cherry picked from commit 9204a9b2c2f6eb59767ab0bddcde62c75e8c20b0)	2020-09-01 09:27:58 +03:00
Yossi Gottlieb	8d79702d8a	Tests: fix redis-cli with remote hosts. (#7693 ) (cherry picked from commit f80f3f492a0ca56e163899eeca7ad40d67d903be)	2020-09-01 09:27:58 +03:00
杨博东	113d5ae872	Fix flock cluster config may cause failure to restart after kill -9 (#7674 ) After fork, the child process(redis-aof-rewrite) will get the fd opened by the parent process(redis), when redis killed by kill -9, it will not graceful exit(call prepareForShutdown()), so redis-aof-rewrite thread may still alive, the fd(lock) will still be held by redis-aof-rewrite thread, and redis restart will fail to get lock, means fail to start. This issue was causing failures in the cluster tests in github actions. Co-authored-by: Oran Agra <oran@redislabs.com> (cherry picked from commit cbaf3c5bbafd43e009a2d6b38dd0e9fc450a3e12)	2020-09-01 09:27:58 +03:00
Yossi Gottlieb	1180537589	Module API: fix missing RM_CLIENTINFO_FLAG_SSL. (#7666 ) The `REDISMODULE_CLIENTINFO_FLAG_SSL` flag was already a part of the `RedisModuleClientInfo` structure but was not implemented. (cherry picked from commit 64c360c5156ca6ee6d1eb52bfeb3fa48f3b25da5)	2020-09-01 09:27:58 +03:00
Oran Agra	916b215fc5	fix new rdb test failing on timing issues (#7604 ) apparenlty on github actions sometimes 500ms is not enough (cherry picked from commit 824bd2ac11472b7a3fce9fcf3189a8e6c6048115)	2020-09-01 09:27:58 +03:00
Oran Agra	a5294c4e52	module hook for master link up missing on successful psync (#7584 ) besides, hooks test was time sensitive. when the replica managed to reconnect quickly after the client kill, the test would fail (cherry picked from commit f7e77759902aa19cfa537ed454e6bc987498e8c5)	2020-09-01 09:27:58 +03:00
WuYunlong	7200b3aa0f	Fix running single test 14-consistency-check.tcl (#7587 ) (cherry picked from commit f3352daf4f4826e1cad4c163fd6e35b81a72e21b)	2020-09-01 09:27:58 +03:00
Yossi Gottlieb	c68b3908a6	Fix TLS cluster tests. (#7578 ) Fix consistency test added in af5167b7f without considering TLS redis-cli configuration. (cherry picked from commit bedf1b21269dfb8afb584bc24585023af8b2a208)	2020-09-01 09:27:58 +03:00
Oran Agra	67750ce3b3	Fix failing tests due to issues with wait_for_log_message (#7572 ) - the test now waits for specific set of log messages rather than wait for timeout looking for just one message. - we don't wanna sample the current length of the log after an action, due to a race, we need to start the search from the line number of the last message we where waiting for. - when attempting to trigger a full sync, use multi-exec to avoid a race where the replica manages to re-connect before we completed the set of actions that should force a full sync. - fix verify_log_message which was broken and unused (cherry picked from commit 109b5ccdcd6e6b8cecdaeb13a246bc49ce7a61f4)	2020-09-01 09:27:58 +03:00
Jiayuan Chen	096285ab64	Add optional tls verification (#7502 ) Adds an `optional` value to the previously boolean `tls-auth-clients` configuration keyword. Co-authored-by: Yossi Gottlieb <yossigo@gmail.com> (cherry picked from commit f31260b0445f5649449da41555e1272a40ae4af7)	2020-09-01 09:27:58 +03:00
Oran Agra	6daa8b9adb	Stabilize bgsave test that sometimes fails with valgrind (#7559 ) on ci.redis.io the test fails a lot, reporting that bgsave didn't end. increaseing the timeout we wait for that bgsave to get aborted. in addition to that, i also verify that it indeed got aborted by checking that the save counter wasn't reset. add another test to verify that a successful bgsave indeed resets the change counter. (cherry picked from commit 8a57969fd75db01b881d438200911d95bdead293)	2020-09-01 09:27:58 +03:00
Oran Agra	e8aa5583d0	testsuite may leave servers alive on error (#7549 ) in cases where you have test name { start_server { start_server { assert } } } the exception will be thrown to the test proc, and the servers are supposed to be killed on the way out. but it seems there was always a bug of not cleaning the server stack, and recently (#7404) we started relying on that stack in order to kill them, so with that bug sometimes we would have tried to kill the same server twice, and leave one alive. luckly, in most cases the pattern is: start_server { test name { } } (cherry picked from commit 36b949438547eb5bf8555fcac2c5040528fd7854)	2020-09-01 09:27:58 +03:00
Yossi Gottlieb	f1d5d5d28e	Tests: drop TCL 8.6 dependency. (#7548 ) This re-implements the redis-cli --pipe test so it no longer depends on a close feature available only in TCL 8.6. Basically what this test does is run redis-cli --pipe, generates a bunch of commands and pipes them through redis-cli, and inspects the result in both Redis and the redis-cli output. To do that, we need to close stdin for redis-cli to indicate we're done so it can flush its buffers and exit. TCL has bi-directional channels can only offers a way to "one-way close" a channel with TCL 8.6. To work around that, we now generate the commands into a file and feed that file to redis-cli directly. As we're writing to an actual file, the number of commands is now reduced. (cherry picked from commit f57e844b2edbb86a5df2f3436045814812c0a3ae)	2020-09-01 09:27:58 +03:00
Remi Collet	af907e4b6d	Fix deprecated tail syntax in tests (#7543 ) (cherry picked from commit 3f2fbc4c614ff718dce7d55fd971d7ed36062c24)	2020-09-01 09:27:58 +03:00
Yossi Gottlieb	b61b663895	Fix oom-score-adj on older distros. (#7724 ) Don't assume `ps` handles `-h` to display output without headers and manually trim headers line from output.	2020-08-30 12:23:47 +03:00
valentinogeron	b7289e912c	EXEC with only read commands should not be rejected when OOM (#7696 ) If the server gets MULTI command followed by only read commands, and right before it gets the EXEC it reaches OOM, the client will get OOM response. So, from now on, it will get OOM response only if there was at least one command that was tagged with `use-memory` flag	2020-08-27 09:19:24 +03:00
Oran Agra	daef1f00c2	Add test coverage for CLIENT UNBLOCK (#7712 ) plus minor other fixes to list.tcl	2020-08-27 08:09:39 +03:00
John Sully	d1f45b8ebf	Implement use-fork config (fails with diskless repl) Former-commit-id: f2d5c2bca22e9fd506db123c47b7f60cdded7e2c	2020-08-24 03:17:59 +00:00
Valentino Geron	9204a9b2c2	Fix LPOS command when RANK is greater than matches When calling to LPOS command when RANK is higher than matches, the return value is non valid response. For example: ``` LPUSH l a :1 LPOS l b RANK 5 COUNT 10 -4 ``` It may break client-side parser. Now, we count how many replies were replied in the array. ``` LPUSH l a :1 LPOS l b RANK 5 COUNT 10 0 ```	2020-08-23 16:03:30 +03:00
Yossi Gottlieb	f80f3f492a	Tests: fix redis-cli with remote hosts. (#7693 )	2020-08-23 10:17:43 +03:00
杨博东	cbaf3c5bba	Fix flock cluster config may cause failure to restart after kill -9 (#7674 ) After fork, the child process(redis-aof-rewrite) will get the fd opened by the parent process(redis), when redis killed by kill -9, it will not graceful exit(call prepareForShutdown()), so redis-aof-rewrite thread may still alive, the fd(lock) will still be held by redis-aof-rewrite thread, and redis restart will fail to get lock, means fail to start. This issue was causing failures in the cluster tests in github actions. Co-authored-by: Oran Agra <oran@redislabs.com>	2020-08-20 08:59:02 +03:00
Yossi Gottlieb	64c360c515	Module API: fix missing RM_CLIENTINFO_FLAG_SSL. (#7666 ) The `REDISMODULE_CLIENTINFO_FLAG_SSL` flag was already a part of the `RedisModuleClientInfo` structure but was not implemented.	2020-08-17 17:46:54 +03:00
Yossi Gottlieb	2530dc0ebd	Add oom-score-adj configuration option to control Linux OOM killer. (#1690 ) Add Linux kernel OOM killer control option. This adds the ability to control the Linux OOM killer oom_score_adj parameter for all Redis processes, depending on the process role (i.e. master, replica, background child). A oom-score-adj global boolean flag control this feature. In addition, specific values can be configured using oom-score-adj-values if additional tuning is required.	2020-08-12 17:58:56 +03:00
Mota	ddcbb628a1	Test:Fix invalid cases in hash.tcl and dump.tcl (#4611 )	2020-08-12 10:25:24 +08:00
Tyson Andre	6f11acbd67	Implement SMISMEMBER key member [member ...] (#7615 ) This is a rebased version of #3078 originally by shaharmor with the following patches by TysonAndre made after rebasing to work with the updated C API: 1. Add 2 more unit tests (wrong argument count error message, integer over 64 bits) 2. Use addReplyArrayLen instead of addReplyMultiBulkLen. 3. Undo changes to src/help.h - for the ZMSCORE PR, I heard those should instead be automatically generated from the redis-doc repo if it gets updated Motivations: - Example use case: Client code to efficiently check if each element of a set of 1000 items is a member of a set of 10 million items. (Similar to reasons for working on #7593) - HMGET and ZMSCORE already exist. This may lead to developers deciding to implement functionality that's best suited to a regular set with a data type of sorted set or hash map instead, for the multi-get support. Currently, multi commands or lua scripting to call sismember multiple times would almost definitely be less efficient than a native smismember for the following reasons: - Need to fetch the set from the string every time instead of reusing the C pointer. - Using pipelining or multi-commands would result in more bytes sent and received by the client for the repeated SISMEMBER KEY sections. - Need to specially encode the data and decode it from the client for lua-based solutions. - Proposed solutions using Lua or SADD/SDIFF could trigger writes to memory, which is undesirable on a redis replica server or when commands get replicated to replicas. Co-Authored-By: Shahar Mor <shahar@peer5.com> Co-Authored-By: Tyson Andre <tysonandre775@hotmail.com>	2020-08-11 11:55:06 +03:00
Meir Shpilraien (Spielrein)	3f494cc49d	see #7544 , added RedisModule_HoldString api. (#7577 ) Added RedisModule_HoldString that either returns a shallow copy of the given String (by increasing the String ref count) or a new deep copy of String in case its not possible to get a shallow copy. Co-authored-by: Itamar Haber <itamar@redislabs.com>	2020-08-09 06:11:47 +03:00
Oran Agra	e2d64485b8	Reduce the probability of failure when start redis in runtest-cluster #7554 (#7635 ) When runtest-cluster, at first, we need to create a cluster use spawn_instance, a port which is not used is choosen, however sometimes we can't run server on the port. possibley due to a race with another process taking it first. such as redis/redis/runs/896537490. It may be due to the machine problem or In order to reduce the probability of failure when start redis in runtest-cluster, we attemp to use another port when find server do not start up. Co-authored-by: Oran Agra <oran@redislabs.com> Co-authored-by: yanhui13 <yanhui13@meituan.com>	2020-08-09 06:08:00 +03:00
Oran Agra	c17e597d05	Accelerate diskless master connections, and general re-connections (#6271 ) Diskless master has some inherent latencies. 1) fork starts with delay from cron rather than immediately 2) replica is put online only after an ACK. but the ACK was sent only once a second. 3) but even if it would arrive immediately, it will not register in case cron didn't yet detect that the fork is done. Besides that, when a replica disconnects, it doesn't immediately attempts to re-connect, it waits for replication cron (one per second). in case it was already online, it may be important to try to re-connect as soon as possible, so that the backlog at the master doesn't vanish. In case it disconnected during rdb transfer, one can argue that it's not very important to re-connect immediately, but this is needed for the "diskless loading short read" test to be able to run 100 iterations in 5 seconds, rather than 3 (waiting for replication cron re-connection) changes in this commit: 1) sync command starts a fork immediately if no sync_delay is configured 2) replica sends REPLCONF ACK when done reading the rdb (rather than on 1s cron) 3) when a replica unexpectedly disconnets, it immediately tries to re-connect rather than waiting 1s 4) when when a child exits, if there is another replica waiting, we spawn a new one right away, instead of waiting for 1s replicationCron. 5) added a call to connectWithMaster from replicationSetMaster. which is called from the REPLICAOF command but also in 3 places in cluster.c, in all of these the connection attempt will now be immediate instead of delayed by 1 second. side note: we can add a call to rdbPipeReadHandler in replconfCommand when getting a REPLCONF ACK from the replica to solve a race where the replica got the entire rdb and EOF marker before we detected that the pipe was closed. in the test i did see this race happens in one about of some 300 runs, but i concluded that this race is unlikely in real life (where the replica is on another host and we're more likely to first detect the pipe was closed. the test runs 100 iterations in 3 seconds, so in some cases it'll take 4 seconds instead (waiting for another REPLCONF ACK). Removing unneeded startBgsaveForReplication from updateSlavesWaitingForBgsave Now that CheckChildrenDone is calling the new replicationStartPendingFork (extracted from serverCron) there's actually no need to call startBgsaveForReplication from updateSlavesWaitingForBgsave anymore, since as soon as updateSlavesWaitingForBgsave returns, CheckChildrenDone is calling replicationStartPendingFork that handles that anyway. The code in updateSlavesWaitingForBgsave had a bug in which it ignored repl-diskless-sync-delay, but removing that code shows that this bug was hiding another bug, which is that the max_idle should have used >= and not >, this one second delay has a big impact on my new test.	2020-08-06 16:53:06 +03:00
WuYunlong	40c7628fc8	Fix tests/cluster/cluster.tcl about wrong usage of lrange. (#6702 )	2020-08-04 18:00:58 +03:00

... 6 7 8 9 10 ...

1910 Commits