futriix

Author	SHA1	Message	Date
antirez	8b15ba933d	Revert "Implements sendfile for redis." This reverts commit 5675053269b0cbc2cf525c99321c96b7c2b39abe.	2020-06-06 11:42:45 +02:00
antirez	3b452cac99	Revert "avoid using sendfile if tls-replication is enabled" This reverts commit 13bbd165e87923558952203d310e9ad053d4d7c0.	2020-06-06 11:42:41 +02:00
antirez	0cd63f06b4	Replication: showLatestBacklog() refactored out.	2020-05-28 10:08:16 +02:00
antirez	e1c7733319	Drop useless line from replicationCacheMaster().	2020-05-27 17:08:51 +02:00
antirez	858845ad56	Remove the meaningful offset feature. After a closer look, the Redis core devleopers all believe that this was too fragile, caused many bugs that we didn't expect and that were very hard to track. Better to find an alternative solution that is simpler.	2020-05-27 12:06:33 +02:00
antirez	7c0fb16790	Merge branch 'unstable' of github.com:/antirez/redis into unstable	2020-05-26 23:55:52 +02:00
antirez	d065e2b7d0	Replication: log backlog creation event.	2020-05-26 23:55:18 +02:00
Salvatore Sanfilippo	34722a4992	Merge pull request #7328 from oranagra/daily_tls_test avoid using sendfile if tls-replication is enabled	2020-05-26 13:19:55 +02:00
Oran Agra	13bbd165e8	avoid using sendfile if tls-replication is enabled this obviously broke the tests, but went unnoticed so far since tls wasn't often tested.	2020-05-26 13:52:06 +03:00
antirez	1bf75eaca7	Clarify what is happening in PR #7320 .	2020-05-25 11:47:38 +02:00
zhaozhao.zz	dd7739f314	PSYNC2: second_replid_offset should be real meaningful offset After adjustMeaningfulReplOffset(), all the other related variable should be updated, including server.second_replid_offset. Or the old version redis like 5.0 may receive wrong data from replication stream, cause redis 5.0 can sync with redis 6.0, but doesn't know meaningful offset.	2020-05-25 11:17:54 +08:00
antirez	7f994abc48	Make disconnectSlaves() synchronous in the base case. Otherwise we run into that: Backtrace: src/redis-server 127.0.0.1:21322(logStackTrace+0x45)[0x479035] src/redis-server 127.0.0.1:21322(sigsegvHandler+0xb9)[0x4797f9] /lib/x86_64-linux-gnu/libpthread.so.0(+0x11390)[0x7fd373c5e390] src/redis-server 127.0.0.1:21322(_serverAssert+0x6a)[0x47660a] src/redis-server 127.0.0.1:21322(freeReplicationBacklog+0x42)[0x451282] src/redis-server 127.0.0.1:21322[0x4552d4] src/redis-server 127.0.0.1:21322[0x4c5593] src/redis-server 127.0.0.1:21322(aeProcessEvents+0x2e6)[0x42e786] src/redis-server 127.0.0.1:21322(aeMain+0x1d)[0x42eb0d] src/redis-server 127.0.0.1:21322(main+0x4c5)[0x42b145] /lib/x86_64-linux-gnu/libc.so.6(__libc_start_main+0xf0)[0x7fd3738a3830] src/redis-server 127.0.0.1:21322(_start+0x29)[0x42b409] Since we disconnect all the replicas and free the replication backlog in certain replication paths, and the code that will free the replication backlog expects that no replica is connected. However we still need to free the replicas asynchronously in certain cases, as documented in the top comment of disconnectSlaves().	2020-05-22 19:29:09 +02:00
antirez	8ecf2998ca	Merge branch 'unstable' of github.com:/antirez/redis into unstable	2020-05-22 16:31:05 +02:00
antirez	ba7f0b7719	Fix #7306 less aggressively. Citing from the issue: btw I suggest we change this fix to something else: * We revert the fix. * We add a call that disconnects chained replicas in the place where we trim the replica (that is a master i this case) offset. This way we can avoid disconnections when there is no trimming of the backlog. Note that we now want to disconnect replicas asynchronously in disconnectSlaves(), because it's in general safer now that we can call it from freeClient(). Otherwise for instance the command: CLIENT KILL TYPE master May crash: clientCommand() starts running the linked of of clients, looking for clients to kill. However it finds the master, kills it calling freeClient(), but this in turn calls replicationCacheMaster() that may also call disconnectSlaves() now. So the linked list iterator of the clientCommand() will no longer be valid.	2020-05-22 16:29:53 +02:00
Salvatore Sanfilippo	42afb7e635	Merge pull request #7096 from ShooterIT/sendfile Implements sendfile for redis.	2020-05-22 13:55:46 +02:00
Salvatore Sanfilippo	51f840fd47	Merge pull request #7305 from madolson/unstable-connection EAGAIN not handled for TLS during diskless load	2020-05-22 12:25:40 +02:00
Qu Chen	7de04d4d36	Disconnect chained replicas when the replica performs PSYNC with the master always to avoid replication offset mismatch between master and chained replicas.	2020-05-21 18:42:10 -07:00
Madelyn Olson	5d6f9cd3b1	EAGAIN for tls during diskless load	2020-05-21 15:20:59 -07:00
antirez	492bbefdb0	Remove the client from CLOSE_ASAP list before caching the master. This was broken in 146201c: we identified a crash in the CI, what was happening before the fix should be like that: 1. The client gets in the async free list. 2. However freeClient() gets called again against the same client which is a master. 3. The client arrived in freeClient() with the CLOSE_ASAP flag set. 4. The master gets cached, but NOT removed from the CLOSE_ASAP linked list. 5. The master client that was cached was immediately removed since it was still in the list. 6. Redis accessed a freed cached master. This is how the crash looked like: === REDIS BUG REPORT START: Cut & paste starting from here === 1092:S 16 May 2020 11:44:09.731 # Redis 999.999.999 crashed by signal: 11 1092:S 16 May 2020 11:44:09.731 # Crashed running the instruction at: 0x447e18 1092:S 16 May 2020 11:44:09.731 # Accessing address: 0xffffffffffffffff 1092:S 16 May 2020 11:44:09.731 # Failed assertion: (:0) ------ STACK TRACE ------ EIP: src/redis-server 127.0.0.1:21300(readQueryFromClient+0x48)[0x447e18] And the 0xffff address access likely comes from accessing an SDS that is set to NULL (we go -1 offset to read the header).	2020-05-16 17:15:35 +02:00
antirez	146201c694	Cache master without checking of deferred close flags. The context is issue #7205: since the introduction of threaded I/O we close clients asynchronously by default from readQueryFromClient(). So we should no longer prevent the caching of the master client, to later PSYNC incrementally, if such flags are set. However we also don't want the master client to be cached with such flags (would be closed immediately after being restored). And yet we want a way to understand if a master was closed because of a protocol error, and in that case prevent the caching.	2020-05-15 10:19:13 +02:00
Oran Agra	5633862924	Keep track of meaningful replication offset in replicas too Now both master and replicas keep track of the last replication offset that contains meaningful data (ignoring the tailing pings), and both trim that tail from the replication backlog, and the offset with which they try to use for psync. the implication is that if someone missed some pings, or even have excessive pings that the promoted replica has, it'll still be able to psync (avoid full sync). the downside (which was already committed) is that replicas running old code may fail to psync, since the promoted replica trims pings form it's backlog. This commit adds a test that reproduces several cases of promotions and demotions with stale and non-stale pings Background: The mearningful offset on the master was added recently to solve a problem were the master is left all alone, injecting PINGs into it's backlog when no one is listening and then gets demoted and tries to replicate from a replica that didn't have any of the PINGs (or at least not the last ones). however, consider this case: master A has two replicas (B and C) replicating directly from it. there's no traffic at all, and also no network issues, just many pings in the tail of the backlog. now B gets promoted, A becomes a replica of B, and C remains a replica of A. when A gets demoted, it trims the pings from its backlog, and successfully replicate from B. however, C is still aware of these PINGs, when it'll disconnect and re-connect to A, it'll ask for something that's not in the backlog anymore (since A trimmed the tail of it's backlog), and be forced to do a full sync (something it didn't have to do before the meaningful offset fix). Besides that, the psync2 test was always failing randomly here and there, it turns out the reason were PINGs. Investigating it shows the following scenario: cycle 1: redis #1 is master, and all the rest are direct replicas of #1 cycle 2: redis #2 is promoted to master, #1 is a replica of #2 and #3 is replica of #1 now we see that when #1 is demoted it prints: 17339:S 21 Apr 2020 11:16:38.523 * Using the meaningful offset 3929963 instead of 3929977 to exclude the final PINGs (14 bytes difference) 17339:S 21 Apr 2020 11:16:39.391 * Trying a partial resynchronization (request e2b3f8817735fdfe5fa4626766daa938b61419e5:3929964). 17339:S 21 Apr 2020 11:16:39.392 * Successful partial resynchronization with master. and when #3 connects to the demoted #2, #2 says: 17339:S 21 Apr 2020 11:16:40.084 * Partial resynchronization not accepted: Requested offset for secondary ID was 3929978, but I can reply up to 3929964 so the issue here is that the meaningful offset feature saved the day for the demoted master (since it needs to sync from a replica that didn't get the last ping), but it didn't help one of the other replicas which did get the last ping.	2020-04-27 15:52:23 +02:00
ShooterIT	5675053269	Implements sendfile for redis.	2020-04-14 23:56:34 +08:00
zhaozhao.zz	9a1be653ad	PSYNC2: reset backlog_idx and master_repl_offset correctly	2020-03-28 20:59:01 +08:00
antirez	94901f61b8	PSYNC2: fix backlog_idx when adjusting for meaningful offset See #7002.	2020-03-27 16:20:02 +01:00
antirez	21976106a9	PSYNC2: meaningful offset implemented. A very commonly signaled operational problem with Redis master-replicas sets is that, once the master becomes unavailable for some reason, especially because of network problems, many times it wont be able to perform a partial resynchronization with the new master, once it rejoins the partition, for the following reason: 1. The master becomes isolated, however it keeps sending PINGs to the replicas. Such PINGs will never be received since the link connection is actually already severed. 2. On the other side, one of the replicas will turn into the new master, setting its secondary replication ID offset to the one of the last command received from the old master: this offset will not include the PINGs sent by the master once the link was already disconnected. 3. When the master rejoins the partion and is turned into a replica, its offset will be too advanced because of the PINGs, so a PSYNC will fail, and a full synchronization will be required. Related to issue #7002 and other discussion we had in the past around this problem.	2020-03-25 15:26:37 +01:00
antirez	8dfbee0f89	Improve comments of replicationCacheMasterUsingMyself().	2020-03-23 16:17:35 +01:00
antirez	ccae1031e7	Merge branch 'unstable' of github.com:/antirez/redis into unstable	2020-03-12 13:25:01 +01:00
Salvatore Sanfilippo	3b68a38464	Merge pull request #6687 from jtru/systemd-integration-fixes Signal systemd readiness atfer Partial Resync	2020-03-06 13:15:10 +01:00
antirez	b3134c8dd3	Make sync RDB deletion configurable. Default to no.	2020-03-04 17:44:21 +01:00
antirez	dff91688a6	Check that the file exists in removeRDBUsedToSyncReplicas().	2020-03-04 12:55:49 +01:00
antirez	994b992782	Log RDB deletion in persistence-less instances.	2020-03-04 11:19:55 +01:00
antirez	b661704fb5	Introduce bg_unlink().	2020-03-04 11:10:54 +01:00
antirez	1d048a3b46	Remove RDB files used for replication in persistence-less instances.	2020-03-03 14:58:15 +01:00
Hengjian Tang	6b55447c1d	modify the read buf size according to the write buf size PROTO_IOBUF_LEN defined before	2020-02-25 15:55:28 +08:00
Salvatore Sanfilippo	8052340ce8	Merge pull request #6822 from guybe7/diskless_load_module_hook_fix Diskless-load emptyDb-related fixes	2020-02-06 13:10:00 +01:00
Guy Benoish	b2bd353f62	Diskless-load emptyDb-related fixes 1. Call emptyDb even in case of diskless-load: We want modules to get the same FLUSHDB event as disk-based replication. 2. Do not fire any module events when flushing the backups array. 3. Delete redundant call to signalFlushedDb (Called from emptyDb).	2020-02-06 16:48:02 +05:30
Salvatore Sanfilippo	cdd502da4c	Merge pull request #6848 from oranagra/opt_use_diskless_load_calls reduce repeated calls to use_diskless_load	2020-02-06 10:30:39 +01:00
Oran Agra	3a74aa3ebe	move restartAOFAfterSYNC from replicaofCommand to replicationUnsetMaster replicationUnsetMaster can be called from other places, not just replicaofCOmmand, and all of these need to restart AOF	2020-02-06 10:14:32 +02:00
Oran Agra	3ea0716cad	reduce repeated calls to use_diskless_load this function possibly iterates on the module list	2020-02-06 09:41:45 +02:00
ShooterIT	fd67d71928	Rename rdb asynchronously	2019-12-31 21:45:32 +08:00
Johannes Truschnigg	eaf450c301	Signal systemd readiness atfer Partial Resync "Partial Resynchronization" is a special variant of replication success that we have to tell systemd about if it is managing redis-server via a Type=Notify service unit.	2019-12-19 21:47:24 +01:00
Johannes Truschnigg	fb25e68f91	Use libsystemd's sd_notify for communicating redis status to systemd Instead of replicating a subset of libsystemd's sd_notify(3) internally, use the dynamic library provided by systemd to communicate with the service manager. When systemd supervision was auto-detected or configured, communicate the actual server status (i.e. "Loading dataset", "Waiting for master<->replica sync") to systemd, instead of declaring readiness right after initializing the server process.	2019-11-19 18:55:44 +02:00
Oran Agra	68c6aacf3b	Modules hooks: complete missing hooks for the initial set of hooks * replication hooks: role change, master link status, replica online/offline * persistence hooks: saving, loading, loading progress * misc hooks: cron loop, shutdown, module loaded/unloaded * change the way hooks test work, and add tests for all of the above startLoading() now gets flag indicating what is loaded. stopLoading() now gets an indication of success or failure. adding startSaving() and stopSaving() with similar args and role.	2019-10-29 17:59:09 +02:00
Wander Hillen	3ca35b8024	Merge branch 'unstable' into minor-typos	2019-10-25 10:18:26 +02:00
Yossi Gottlieb	85d7f38136	Merge remote-tracking branch 'upstream/unstable' into tls	2019-10-16 17:08:07 +03:00
Salvatore Sanfilippo	15e422686b	Merge pull request #6429 from charsyam/feature/typo-slave [trivial] fix typos salves to slaves in replication.c	2019-10-10 14:56:43 +02:00
antirez	9aba7564d6	Cluster: fix memory leak of cached master. This is what happened: 1. Instance starts, is a slave in the cluster configuration, but actually server.masterhost is not set, so technically the instance is acting like a master. 2. loadDataFromDisk() calls replicationCacheMasterUsingMyself() even if the instance is a master, in the case it is logically a slave and the cluster is enabled. So now we have a cached master even if the instance is practically configured as a master (from the POV of server.masterhost value and so forth). 3. clusterCron() sees that the instance requires to replicate from its master, because logically it is a slave, so it calls replicationSetMaster() that will in turn call replicationCacheMasterUsingMyself(): before this commit, this call would overwrite the old cached master, creating a memory leak.	2019-10-10 10:23:34 +02:00
Yossi Gottlieb	df08b624bd	TLS: Configuration options. Add configuration options for TLS protocol versions, ciphers/cipher suites selection, etc.	2019-10-07 21:07:27 +03:00
Oran Agra	6fd5ff8c98	diskless replication rdb transfer uses pipe, and writes to sockets form the parent process. misc: - handle SSL_has_pending by iterating though these in beforeSleep, and setting timeout of 0 to aeProcessEvents - fix issue with epoll signaling EPOLLHUP and EPOLLERR only to the write handlers. (needed to detect the rdb pipe was closed) - add key-load-delay config for testing - trim connShutdown which is no longer needed - rioFdsetWrite -> rioFdWrite - simplified since there's no longer need to write to multiple FDs - don't detect rdb child exited (don't call wait3) until we detect the pipe is closed - Cleanup bad optimization from rio.c, add another one	2019-10-07 21:06:30 +03:00
Yossi Gottlieb	10ffeb03e4	TLS: Connections refactoring and TLS support. * Introduce a connection abstraction layer for all socket operations and integrate it across the code base. * Provide an optional TLS connections implementation based on OpenSSL. * Pull a newer version of hiredis with TLS support. * Tests, redis-cli updates for TLS support.	2019-10-07 21:06:13 +03:00

1 2 3 4 5 ...

311 Commits