perf(stream): trim with a single DeleteRange instead of per-entry Delete (#3592)

### Problem

`Stream::trim()` — the shared path behind `XTRIM` and `XADD` with
`MAXLEN`/`MINID` — removes trimmed entries one at a time: a per-entry
`WriteBatch::Delete()` in a loop, committed as a single atomic batch.
This makes trimming **O(entries removed)**:

- Removing a large number of entries builds a correspondingly large
write batch of point tombstones and commits it in one atomic write.
- Those point tombstones are then scanned one-by-one by every reader,
and re-scanned by each subsequent trim's `Seek`.

For a stream that is trimmed continuously (each trim removes many
entries), this is expensive both on the write and on subsequent
reads/trims.

### Change

Replace the per-entry `Delete()` loop with a single
`WriteBatch::DeleteRange()` over the trimmed span — the same approach
already used for bulk deletion in `DestroyGroup()`.

- The entries are still scanned to count them, so `metadata` (`size`,
`first_entry_id`, `max_deleted_entry_id`) is maintained exactly as
before and `XLEN` stays exact. Only the deletion mechanism changes: the
write batch stays O(1) (one range tombstone) instead of O(entries
removed).
- Group, consumer and PEL subkeys are encoded with a `UINT64_MAX` prefix
and sort after every entry key, so they always stay outside the deleted
range. The exclusive upper bound is the first surviving entry for a
partial trim, or the next-version prefix when the whole stream is
trimmed.

### Tests

Adds cases for the behaviour specific to the range delete, alongside the
existing trim suite:

- consumer group + PEL survival across a full-entry trim,
- entries added after a trim-to-empty remaining visible (range-tombstone
sequence-number semantics),
- repeated `MINID` trims interleaved with adds keeping exact survivors
and first/max-deleted bookkeeping.

---------

Co-authored-by: hulk <hulk.website@gmail.com>
10 files changed
tree: 0a407620ffb193bfa35b425cc570446be16cf43a
  1. .devcontainer/
  2. .github/
  3. assets/
  4. cmake/
  5. dev/
  6. licenses/
  7. src/
  8. tests/
  9. utils/
  10. .asf.yaml
  11. .clang-format
  12. .clang-tidy
  13. .dockerignore
  14. .gitignore
  15. AGENTS.md
  16. CMakeLists.txt
  17. Dockerfile
  18. kvrocks.conf
  19. LICENSE
  20. NOTICE
  21. README.md
  22. SECURITY.md
  23. THREAT_MODEL.md
  24. x.py
README.md

CI License GitHub stars


Apache Kvrocks is a distributed key value NoSQL database that uses RocksDB as storage engine and is compatible with Redis protocol. Kvrocks intends to decrease the cost of memory and increase the capacity while compared to Redis. The design of replication and storage was inspired by rocksplicator and blackwidow.

Kvrocks has the following key features:

  • Redis Compatible: Users can access Apache Kvrocks via any Redis client.
  • Namespace: Similar to Redis SELECT but equipped with token per namespace.
  • Replication: Async replication using binlog like MySQL.
  • High Availability: Support Redis sentinel to failover when master or slave was failed.
  • Cluster: Centralized management but accessible via any Redis cluster client.

Who uses Kvrocks

You can find Kvrocks users at the Users page.

Users are encouraged to add themselves to the Users page. Either leave a comment on the “Who is using Kvrocks” issue, or directly send a pull request to add company or organization information and logo.

Build and run Kvrocks

Prerequisite

# Ubuntu / Debian
sudo apt update
sudo apt install -y git build-essential cmake libtool python3 libssl-dev

# CentOS / RedHat
sudo yum install -y centos-release-scl-rh
sudo yum install -y git devtoolset-11 autoconf automake libtool libstdc++-static python3 openssl-devel
# download and install cmake via https://cmake.org/download
wget https://github.com/Kitware/CMake/releases/download/v3.26.4/cmake-3.26.4-linux-x86_64.sh -O cmake.sh
sudo bash cmake.sh --skip-license --prefix=/usr
# enable gcc and make in devtoolset-11
source /opt/rh/devtoolset-11/enable

# openSUSE / SUSE Linux Enterprise
sudo zypper install -y gcc11 gcc11-c++ make wget git autoconf automake python3 curl cmake

# Arch Linux
sudo pacman -Sy --noconfirm autoconf automake python3 git wget which cmake make gcc

# macOS
brew install git cmake autoconf automake libtool openssl
# please link openssl by force if it still cannot be found after installing
brew link --force openssl

Build

It is as simple as:

$ git clone https://github.com/apache/kvrocks.git
$ cd kvrocks
$ ./x.py build # `./x.py build -h` to check more options

To build with TLS support, you'll need OpenSSL development libraries (e.g. libssl-dev on Debian/Ubuntu) and run:

$ ./x.py build -DENABLE_OPENSSL=ON

To build with lua instead of luaJIT, run:

$ ./x.py build -DENABLE_LUAJIT=OFF

Build with debug mode, run:

# The default build type is RelWithDebInfo and its optimization level is typically -O2.
# You can change it to -O0 in debug mode.

$ ./x.py build -DCMAKE_BUILD_TYPE=Debug

Running Kvrocks

$ ./build/kvrocks -c kvrocks.conf

Running Kvrocks using Docker

$ docker run -it -p 6666:6666 apache/kvrocks --bind 0.0.0.0
# or get the nightly image:
$ docker run -it -p 6666:6666 apache/kvrocks:nightly

Please visit Apache Kvrocks on DockerHub for additional details about images.

Connect Kvrocks service

$ redis-cli -p 6666

127.0.0.1:6666> get a
(nil)

Running test cases

$ ./x.py build --unittest
$ ./x.py test cpp # run C++ unit tests
$ ./x.py test go # run Golang (unit and integration) test cases

Supported platforms

  • OS: Linux and macOS
  • arch: x86_64, ARM and RISC-V

Namespace

Namespace is used to isolate data between users. Unlike all the Redis databases can be visited by requirepass, we use one token per namespace. requirepass is regarded as admin token, and only admin token allows to access the namespace command, as well as some commands like config, slaveof, bgsave, etc. See the Namespace page for more details.

# add token
127.0.0.1:6666> namespace add ns1 my_token
OK

# update token
127.0.0.1:6666> namespace set ns1 new_token
OK

# list namespace
127.0.0.1:6666> namespace get *
1) "ns1"
2) "new_token"
3) "__namespace"
4) "foobared"

# delete namespace
127.0.0.1:6666> namespace del ns1
OK

Cluster

Kvrocks implements a proxyless centralized cluster solution but its accessing method is completely compatible with Redis cluster clients. You can use Redis cluster SDKs to access the kvrocks cluster. For more details, please refer to Kvrocks Cluster Introduction.

Documents

Documents are hosted at the official website.

Tools

  • To manage Kvrocks clusters for failover, scaling up/down and more, use kvrocks-controller
  • To export the Kvrocks monitor metrics, use kvrocks_exporter
  • To migrate from Redis to Kvrocks, use RedisShake
  • To migrate from Kvrocks to Redis, use kvrocks2redis built via ./x.py build

Contributing

Kvrocks community welcomes all forms of contribution and you can find out how to get involved on the Community and How to Contribute pages.

License

Apache Kvrocks is licensed under the Apache License Version 2.0. See LICENSE and NOTICE for details.

Social Media

WeChat official account