diff --git a/.agents/skills/automate-test/SKILL.md b/.agents/skills/automate-test/SKILL.md new file mode 100644 index 000000000..8cec41058 --- /dev/null +++ b/.agents/skills/automate-test/SKILL.md @@ -0,0 +1,110 @@ +--- +name: automate-test +description: 开发环境测试闭环:单元测试 / e2e / 重型回归 / xfstests。 +context: fork +disable-model-invocation: true +--- + +# dingofs 自动化测试技能 + +自动化测试闭环:跑测试 → 定位失败 → 修代码 → 重编译重部署 → 再跑。仅限开发环境,不用于生产,不自行 git 提交。 + +## 前置:服务就绪 + +编译、MDS/client 的部署启动、日志位置全部见 `/skill:dev-deploy`。本技能假设 dist/ 下服务已在跑,**不重复部署步骤**。 + +## 编译 + +```bash +cd build && make -j 12 +``` + +单元测试二进制要求 build 以 `-DBUILD_UNIT_TESTS=ON` 配置(当前 build/ 已满足)。若 `build/bin/test` 不存在或为空,重新配置: + +```bash +cd build && cmake -DCMAKE_BUILD_TYPE=RelWithDebInfo -DBUILD_UNIT_TESTS=ON .. && make -j 12 +``` + +## 单元测试 + +`build/bin/test` 下是 gtest 二进制,逐个直接运行,无需参数。`test_coverage_helper` 是覆盖率辅助程序,跳过。 + +判据:进程退出码为 0,且输出中无 `[ FAILED ]`。 + +## 端到端与重型工具 + +`run_all_test.sh` 封装了 e2e / pjdfstest / fsx / mdtest / fio / fsstress。**`--mountpoint` 必填**,不传直接退出。 + +```bash +cd scripts/dev-mds +bash run_all_test.sh --mountpoint=$MOUNT_POINT --type=e2e --round=1 +``` + +| 场景 | `--type` | +|---|---| +| 日常回归 | `e2e` | +| 全量重型(e2e / pjdfstest / fsx / mdtest / fsstress) | `all` | +| 单个工具 | `pjdtest` \| `fsx` \| `mdtest` \| `fio` \| `fsstress` | + +- `--round` 默认 1;只有反复跑找偶发才需要调大。 +- e2e 依赖 `test/e2e` 的 uv 环境;pjdfstest 依赖 `/home/dengzihui/work/dingofs-test/pjdfstest/tests` 存在。 +- `--mds-addr` 在脚本里已定义但未使用,不要传。 + +判据:退出码 0,且输出中没有 `result: FAIL` —— **脚本只对 e2e / pjdfstest / fsx 判定**;mdtest / fio / fsstress 不判定,需自己查日志确认无 error。 + +日志:`/tmp/dev-regression-test/_<时间戳>_<轮次>/`。 + +## xfstests(脚本未覆盖) + +适配层的安装、产物、local/MDS 两种模式见仓库 `xfstests/README_CN.md`,不在此重复。 + +```bash +# 一次性:安装挂载 helper 并生成配置;meta-url 按需改 +DINGOFS_META_URL_TEMPLATE="mds://:7801/{fsname}" \ + bash xfstests/setup.sh /home/dengzihui/work/dingofs-test/xfstests-dev +``` + +```bash +cd /home/dengzihui/work/dingofs-test/xfstests-dev +sudo ./check $(cat tests/generic/supported) +``` + +`` 从 `scripts/dev-mds/mds_deploy_parameters.local` 取。 + +判据:`./check` 退出码 0(末行 `Passed all ... tests`)。 +失败证据:`results/generic/NNN.out.bad`(测试侧)、`/mnt/dingofs-xfstests/runtime//log/`(client 侧)。 + +## vdbench(脚本未覆盖) + +```bash +cd /home/dengzihui/work/dingofs-test/vdbench +# 先改 config/test-01.vd:anchor= 指到 $MOUNT_POINT 下,elapsed 改成回归可接受的秒数 +./vdbench -f config/test-01.vd +``` + +判据:退出码 0 且输出无 error。结果在 `output/logfile.html`、`output/flatfile.html`。 + +## 测试对象地址 + +第 1 个 MDS 实例监听 `:`(默认 7801),取值见 `scripts/dev-mds/mds_deploy_parameters.local`。 + +- meta: `mds://:7801/` +- fs 不存在时先创建(需要 MDS 已在跑):`cd scripts/dev-mds && bash create_fs.sh --fs_name=$FS_NAME --mds_addr=:7801` + +## 流程 + +1. **编译**:`cd build && make -j 12` 成功。 +2. **确认服务在跑**:`pgrep -c -x dingo-mds` 等于 `SERVER_NUM`,且 `mountpoint -q $MOUNT_POINT` 成立;不一致先按 `/skill:dev-deploy` 重部署。(不要用 `ps -ef | grep`:会匹配到自己,且本机可能另有部署的 client。) +3. **执行测试**:单元测试逐个跑,或 `run_all_test.sh --type=...`;需要时加 xfstests / vdbench。记录到 trace(见下)。 +4. **判定**:按各自判据全绿 → 跳 6;有失败 → 下一步。 +5. **定位并修复**:e2e 失败看对应 `result` 上方的日志目录,单元测试看 stderr,结合 `dist/*/log/` 里的服务日志定位根因,改代码后回到 1。 + **同一个测试连续 3 轮仍不通过就停手**,把已定位的根因、试过的改法、日志路径汇报给用户,不要继续盲改。 +6. **报告**:给出变更清单与测试结论。**不要自行 git 提交**,提交交给用户或 `/skill:git-commit`。 + +## 跟踪 + +每完成上面一步,向 `/tmp/automate-test.trace` 追加一行: + +``` +<时间> <步骤号> <命令> <结果> <日志路径> +``` diff --git a/.agents/skills/dev-deploy/SKILL.md b/.agents/skills/dev-deploy/SKILL.md index 7daf3adf0..d5c84b03f 100644 --- a/.agents/skills/dev-deploy/SKILL.md +++ b/.agents/skills/dev-deploy/SKILL.md @@ -1,73 +1,98 @@ --- name: dev-deploy -description: 在开发环境部署dingofs的技能,当开发完成功能或修复bug后,可以使用这个技能将代码部署到测试环境进行验证 +description: 开发环境部署/重部署 dingofs:部署并启动 MDS 与 client,再做基本可用性冒烟。开发完成功能或修复 bug 后,需要把改动跑起来(重启服务)验证时使用。仅限开发环境,不做性能/压力/稳定性测试。 context: fork -disable-model-invocation: true --- +# dingofs 部署技能 -# dingofs部署技能 -**注意**: 本技能仅适用于开发环境部署测试dingofs,不能用于生产环境部署,并只进行基本功能验证,不用于性能测试或压力测试或者稳定性测试。 - -## 脚本 -脚本在`scripts/dev-mds`目录下,包含以下文件: -- ** mds_deploy_parameters **: 部署mds服务器的参数配置文件,包含服务器实例数量、起始端口号、起始实例ID等参数。 - - CLUSTER_ID: 集群ID,默认为101 - - SERVER_NUM: 服务器实例数量,默认为1 - - SERVER_HOST: 服务器主机IP地址 - - SERVER_LISTEN_HOST: 服务器监听IP地址 - - SERVER_START_PORT: 服务器起始端口号,默认为7800 - - MDS_INSTANCE_START_ID: mds服务器实例起始ID,默认为1000 - - COORDINATOR_ADDR: coordinator地址 -- ** deploy_mds.sh **: 部署mds服务器的脚本,支持部署多个实例,并且可以选择是否替换配置文件。 -- ** start_mds.sh **: 启动mds服务器的脚本,支持启动多个实例,并且可以选择是否替换配置文件。 -- ** stop_mds.sh **: 停止mds服务器的脚本,支持停止多个实例。 -- ** start_client.sh **: 部署和启动client的脚本 - -## MDS部署、启动、停止 -**注意**: 必须在脚本目录scripts/dev-mds下执行以下命令,否则会报错。 -部署目标目录为项目根目录下的dist目录。 -```bash -# 进入脚本目录 -cd scripts/dev-mds +仅适用于开发环境:部署 + 基本功能验证。不用于生产环境,也不用于性能/压力/稳定性测试。 -# 部署,部署目录为 dist/mds-1 dist/mds-2 ... dist/mds-N,N为服务器实例数量 -# mds目录下包括:bin、conf、log等目录 -bash deploy_mds.sh --server_num=${SERVER_NUM} +## 参数 -# 启动 -bash start_mds.sh --server_num=${SERVER_NUM} +`scripts/dev-mds/mds_deploy_parameters.local` —— 本机私有,已被 `.gitignore` 忽略;入库的模板是 `mds_deploy_parameters`,首次从模板复制。 -# 停止 -bash stop_mds.sh --server_num=${SERVER_NUM} +脚本真正消费的键: -# 一键部署、启动 -bash clean_start.sh --server_num=${SERVER_NUM} +- `SERVER_NUM`:MDS 实例数 +- `SERVER_HOST` / `SERVER_LISTEN_HOST`:对外 / 监听地址 +- `SERVER_START_PORT`:起始端口;第 `i` 个实例为 `SERVER_START_PORT + i`,即首实例 `7801` +- `CLUSTER_ID`、`MDS_INSTANCE_START_ID`、`COORDINATOR_ADDR` +- `STORAGE_ENGINE` / `STORAGE_URL`:由 `deploy_mds.sh` 写进 `mds.conf` +- `S3_ENDPOINT` / `S3_AK` / `S3_SK` / `S3_BUCKETNAME`、`LOCAL_DATASTORE_PATH`:仅 `create_fs.sh` 使用 -``` +## 步骤 +以下命令**必须**在 `scripts/dev-mds` 目录下执行,否则脚本会报错。部署产物在项目根目录 `dist/`。 -## Client部署启动 -**注意**: 必须在脚本目录scripts/dev-mds下执行以下命令,否则会报错。 -```bash +1. **编译** + + ```bash + cd build && make -j 12 + ``` + + 判据:退出码 0。 + +2. **停止 client** + + 先停 client 再动 MDS —— MDS 的二进制会被重新软链,挂着旧进程容易踩坑。 + + ```bash + cd scripts/dev-mds + sudo ./start_client.sh --meta=$META_ADDR --mountpoint=$MOUNT_POINT --num=1 --stop + ``` + +3. **部署并启动 MDS** + + ```bash + bash clean_start.sh --server_num=$SERVER_NUM + ``` + + 判据:`pgrep -c dingo-mds` 输出等于 `SERVER_NUM`。(不要用 `ps -ef | grep`,它会匹配到自己,也不校验数量。) -# 进入脚本目录 -cd scripts/dev-mds + 失败看 `dist/mds-/log/out`。 -# 部署和启动client (需要先部署启动mds服务器) -# META_ADDR: mds服务器的地址,格式为mds://ip:port/fs_name 例如mds://10.220.69.5:7801/dengzh_hash_01 +4. **启动 client** + `META_ADDR` 格式为 `mds://:/`,例如 `mds://10.220.69.5:7801/dengzh_hash_01`。fs 不存在时先创建: -sudo ./start_client.sh --meta=$META_ADDR --mountpoint=$MOUNT_POINT --num=1 --noupgrade --clean_log + ```bash + bash create_fs.sh --fs_name=$FS_NAME --mds_addr=$SERVER_HOST:$(($SERVER_START_PORT + 1)) + ``` + ```bash + sudo ./start_client.sh --meta=$META_ADDR --mountpoint=$MOUNT_POINT \ + --num=1 --noupgrade --clean_log + ``` -# 停止 -sudo ./start_client.sh --meta=$META_ADDR --mountpoint=$MOUNT_POINT --num=1 --stop + - `--noupgrade`:不重装 client 二进制,直接用 `dist/client/bin` 下已有的 + - `--clean_log`:清掉 `dist/client/log/` 的旧日志(排查问题时建议保留,去掉此参数) + 判据:`mountpoint -q $MOUNT_POINT`。 + +5. **冒烟验证** + + ```bash + touch $MOUNT_POINT/.deploy_check && rm $MOUNT_POINT/.deploy_check + ``` + + 判据:退出码 0。失败先看 `dist/client/log/` 与 `dist/mds-*/log/`。 + +6. **汇报** + + 给用户:MDS 实例数、client 挂载点、冒烟结果、日志路径。**不要自行 git 提交**,提交交给用户或 `/skill:git-commit`。 + +需要跑测试(单元 / e2e / 重型回归 / xfstests)→ `/skill:automate-test`。 + +## 首次环境 + +集群和文件系统尚不存在时,MDS 起来之后先建(`create_cluster.sh` 的 `--cluster_id` 必须大于 0): + +```bash +bash create_cluster.sh --cluster_id=101 +bash create_fs.sh --fs_name=$FS_NAME --mds_addr=$SERVER_HOST:$(($SERVER_START_PORT + 1)) ``` +## 其他脚本 -## 步骤 -1. 代码变更后,重新编译代码,编译成功后进入后续步骤。 -2. 一键部署、启动mds服务器,使用clean_start.sh脚本,启动成功后进入后续步骤,可以使用`ps -ef | grep dingo-mds`命令查看mds服务器是否启动成功。 -3. 先停止client,再启动client,使用start_client.sh脚本,启动成功后进入后续步骤,可以使用`ps -ef | grep dingo-client`命令查看client是否启动成功。 +`deploy_mds.sh` / `start_mds.sh` / `stop_mds.sh` 是 `clean_start.sh` 的拆分,只重启 MDS 时单独用。参数以 `bash <脚本> --help` 为准。 diff --git a/.agents/skills/dev-regression-test/SKILL.md b/.agents/skills/dev-regression-test/SKILL.md deleted file mode 100644 index 52531d01c..000000000 --- a/.agents/skills/dev-regression-test/SKILL.md +++ /dev/null @@ -1,294 +0,0 @@ ---- -name: dev-regression-test -description: 在开发环境进行回归测试,当开发完成功能或修复bug后,可以使用这个技能在开发环境进行回归测试,验证代码变更是否生效,是否引入新问题。 -context: fork -disable-model-invocation: true ---- - - -# dingofs回归测试技能 -**注意**: 本技能仅适用于开发环境部署测试dingofs,不要用于生产环境部署,并只进行基本功能回归验证,提交代码之前可以使用这个技能在开发环境进行回归测试,验证代码变更是否生效,是否引入新问题。 - -使用方式: /dev-regression-test [测试目录路径] - - -## 测试环境 -服务(dingo-client、dingo-mds)运行在项目目录dist下面,当出现测试问题可以查看对应日志。 - -目录scripts/dev-mds下面是启动和停止服务的脚本,使用方法可以参考脚本中的注释说明。 - -```bash - -# 目录结构 -dengzihui@dingofs-5 ➜ dingofs git:(feat/main_081001) tree dist -dist -├── cache -│   ├── bin -│   │   └── dingo-cache -> /home/dengzihui/work/dingofs/build/bin/dingo-cache -│   └── log -├── client -│   ├── bin -│   │   └── dingo-client -> /home/dengzihui/work/dingofs/build/bin/dingo-client -│   ├── cache -│   │   └── dengzh_hash_01-1 -│   ├── conf -│   └── log -├── mds-1 -│   ├── bin -│   │   ├── dingo-mds -> /home/dengzihui/work/dingofs/build/bin/dingo-mds -│   │   └── dingo-mds-client -> /home/dengzihui/work/dingofs/build/bin/dingo-mds-client -│   ├── conf -│   │   ├── coor_list -│   │   └── mds.conf -│   └── log -├── mds-2 -│   ├── bin -│   │   ├── dingo-mds -> /home/dengzihui/work/dingofs/build/bin/dingo-mds -│   │   └── dingo-mds-client -> /home/dengzihui/work/dingofs/build/bin/dingo-mds-client -│   ├── conf -│   │   ├── coor_list -│   │   └── mds.conf -│   └── log -└── mds-3 - ├── bin - │   ├── dingo-mds -> /home/dengzihui/work/dingofs/build/bin/dingo-mds - │   └── dingo-mds-client -> /home/dengzihui/work/dingofs/build/bin/dingo-mds-client - ├── conf - │   ├── coor_list - │   └── mds.conf - └── log - -``` - - -## 回归测试工具 - - -### 基础命令 -基础的文件系统操作命令,主要用于测试基本的文件系统功能是否正常,包括创建文件、删除文件、重命名文件、创建目录、删除目录、重命名目录等操作,可以自由使用下面的工具进行测试,根据命令执行结果可以判断基本的文件系统功能是否正常。 -测试方法: 使用下面的工具进行基本的文件系统操作测试,参数可以自己指定。 -```bash -# 下面的工具可以使用--help参数查看具体用法和参数说明 - -# 创建目录 -mkdir -# 删除目录 -rm -# 创建文件 -touch -# 删除文件 -rm -# 重命名文件 -mv -# 列出目录内容 -ls -# 写入数据到文件 -dd -# 显示文件内容 -cat -# 显示文件属性 -stat -# 修改文件权限 -chmod -# 修改文件所有者 -chown -# 截断文件 -truncate - -``` - -### pjdtest工具 -主要用于测试元数据操作是否正常,是否符合 POXIS 标准,测试内容包括创建文件、删除文件、重命名文件、创建目录、删除目录、重命名目录等操作。 -测试方法: 使用prove命令运行pjdfstest目录下的测试用例,参数可以自己指定,注意必须用sudo权限。 -```bash - -# 环境信息 -PJD_SUFFIX=$(date +%Y%m%d%H%M%S) -PJD_TEST_DIR=$ARGUMENTS[0]/pjd_test_${PJD_SUFFIX} -PJD_LOG_DIR=/tmp/dev-regression-test/pjd_test_${PJD_SUFFIX} - - - -# 创建测试目录和日志目录 -mkdir -p ${PJD_TEST_DIR} -mkdir -p ${PJD_LOG_DIR} - -# 必须跳转到测试目录 -cd ${PJD_TEST_DIR} - -# 注意:必须用sudo权限运行测试用例,否则会出现权限问题,导致测试失败。 - -# 示例: 运行全部测试用例 -sudo prove -rv --exec 'bash -x' /home/dengzihui/work/dingofs-test/pjdfstest/tests > $PJD_LOG_DIR/pjd_test.log 2>&1 - -# 示例: 运行部分测试用例 -sudo prove -rv --exec 'bash -x' /home/dengzihui/work/dingofs-test/pjdfstest/tests/mknod > $PJD_LOG_DIR/pjd_test.log 2>&1 - -# 示例: 运行一个测试用例 -sudo prove -rv --exec 'bash -x' /home/dengzihui/work/dingofs-test/pjdfstest/tests/mknod/00.t > $PJD_LOG_DIR/pjd_test.log 2>&1 - - -``` - -### fsx工具 - -```bash - -# 环境信息 -FSX_SUFFIX=$(date +%Y%m%d%H%M%S) -FSX_TEST_FILE=$ARGUMENTS[0]/fsx_test_${FSX_SUFFIX} -FSX_LOG_DIR=/tmp/dev-regression-test/fsx_test_${FSX_SUFFIX} - - -# 创建测试目录和日志目录 -mkdir -p ${FSX_TEST_FILE} -mkdir -p ${FSX_LOG_DIR} - - -# 示例: 运行测试命令 -fsx -l 1073741824 -o 1048576 -S 0 -p 10000 --duration=3600 --record-ops=$LOG_DIR/fsx.ops -P $FSX_LOG_DIR $FSX_TEST_FILE - -``` - -### mdtest工具 -主要用于测试元数据性能,测试内容包括创建文件、删除文件、重命名文件、创建目录、删除目录、重命名目录等操作,测试结果会输出到日志目录下,可以查看日志分析测试结果。 -测试方法: 使用mpirun和mdtest工具结合起来测试,参数可以自己指定。 -```bash - -# 环境信息 -MDTEST_SUFFIX=$(date +%Y%m%d%H%M%S) -MDTEST_TEST_DIR=$ARGUMENTS[0]/mdtest_test_${MDTEST_SUFFIX} -MDTEST_LOG_DIR=/tmp/dev-regression-test/mdtest_test_${MDTEST_SUFFIX} - -# 创建测试目录和日志目录 -mkdir -p ${MDTEST_TEST_DIR} -mkdir -p ${MDTEST_LOG_DIR} - -# 示例: 运行测试命令 -mpirun -np 4 mdtest -z 0 -b 1 -n 1000 -L -C -F -d ${MDTEST_TEST_DIR} > ${MDTEST_LOG_DIR}/mdtest.log 2>&1 - -``` - - -### vdbench工具 -主要用户测试数据读写性能,测试内容包括顺序读写和随机读写,测试结果会输出到日志目录下,可以查看日志分析测试结果。 -测试方法: 使用vdbench工具进行测试,参数可以自己指定。 -```bash - -# 环境信息 -VDBENCH_SUFFIX=$(date +%Y%m%d%H%M%S) -VDBENCH_TEST_DIR=$ARGUMENTS[0]/vdbench_test_${VDBENCH_SUFFIX} -VDBENCH_TOOL_DIR=/home/dengzihui/work/dingofs-test/vdbench -VDBENCH_LOG_DIR=/tmp/dev-regression-test/vdbench_test_${VDBENCH_SUFFIX} - - -# 创建测试目录和日志目录 -mkdir -p ${VDBENCH_TEST_DIR} -mkdir -p ${VDBENCH_LOG_DIR} - - - -# 跳转到工具目录 -cd ${VDBENCH_TOOL_DIR} - -# 注意: ${VDBENCH_TOOL_DIR}/config目录下有一些测试配置,可以根据需要修改配置文件,或者自己创建新的配置文件进行测试,特别注意要修改配置中的目标测试目录 - -# 示例: 运行测试命令 -./vdbench -f config/test-01.vd > ${VDBENCH_LOG_DIR}/vdbench.log 2>&1 - - -``` - - -### fio工具 -主要用于测试数据读写性能,测试内容包括顺序读写和随机读写,测试结果会输出到日志目录下,可以查看日志分析测试结果。 -测试方法: 使用fio工具进行测试,参数可以自己指定。 -```bash - -# 环境信息 -FIO_SUFFIX=$(date +%Y%m%d%H%M%S) -FIO_TEST_DIR=$ARGUMENTS[0]/fio_test_${FIO_SUFFIX} -FIO_LOG_DIR=/tmp/dev-regression-test/fio_test_${FIO_SUFFIX} - - -# 创建测试目录 -mkdir -p ${FIO_TEST_DIR} -mkdir -p ${FIO_LOG_DIR} - -# 跳转到测试目录 -cd ${FIO_TEST_DIR} - -# 示例: 运行测试命令 -fio --ioengine=libaio --iodepth=1 --direct=1 --rw=read --bs=128KB --size=1GB --numjobs=8 --group_reporting --name=test > ${FIO_LOG_DIR}/fio.log 2>&1 - -``` - -### fsstress工具 -主要进行文件系统压力测试和并发测试 -测试方法: 使用fsstress工具进行测试,参数可以自己指定。 -```bash - -# 环境信息 -FSSTRESS_SUFFIX=$(date +%Y%m%d%H%M%S) -FSSTRESS_TEST_DIR=$ARGUMENTS[0]/fsstress_test_${FSSTRESS_SUFFIX} -FSSTRESS_LOG_DIR=/tmp/dev-regression-test/fsstress_test_${FSSTRESS_SUFFIX} - -# 创建测试目录和日志目录 -mkdir -p ${FSSTRESS_TEST_DIR} -mkdir -p ${FSSTRESS_LOG_DIR} - -# 跳转到测试目录 -cd ${FSSTRESS_TEST_DIR} - -# 示例: 运行测试命令 -/opt/ltp/testcases/bin/fsstress -d ${FSSTRESS_TEST_DIR} -n 10000 -p 8 -v > ${FSSTRESS_LOG_DIR}/fsstress.log 2>&1 - -``` - -### xfstests工具 -主要用于测试文件系统语义。 -测试方法: 项目下有xfstests目录,里面有xfstests说明信息,可以参考。 -注意: xfstests测试使用xftest和xfscratch两个文件系统进行测试,不要使用$ARGUMENTS[0]。 -```bash - -# 环境信息 -# xfstests项目目录 -XFSTESTS_TOOL_DIR=/home/dengzihui/work/dingofs-test/xfstests-dev -# xfstests测试运行的根目录,里面包括运行时信息(挂载点、日志等) -XFSTESTS_ROOT_DIR=/mnt/dingofs-xfstests - -# MDS地址 -MDS_ADDR=10.220.69.5:7801 -DINGOFS_META_URL_TEMPLATE=mds://${MDS_ADDR}/{fsname} - -# 创建日志目录 -mkdir -p ${XFSTESTS_LOG_DIR} - -# 初始化环境完成后清理老日志 -rm -rf ${XFSTESTS_ROOT_DIR}/runtime/xftest/log/* -rm -rf ${XFSTESTS_ROOT_DIR}/runtime/xfscratch/log/* - - -# 检查文件系统xftest和xfscratch是否存在,使用dingo-mds-client工具 -# 示例: 检查文件系统xftest是否存在 -dingo-mds-client --cmd=getfs --mds_addr=${MDS_ADDR} --fs_name=xftest -dingo-mds-client --cmd=getfs --mds_addr=${MDS_ADDR} --fs_name=xfscratch - - -# 运行具体测试用例前需要调用脚本xfstests/setup.sh初始化环境,示例: -./xfstests/setup.sh ${XFSTESTS_TOOL_DIR} ${DINGOFS_META_URL_TEMPLATE} - -# 准备好环境后,运行测试用例,只运行generic下面的已支持的测试用例 -# 已支持的测试用例文件: $XFSTESTS_TOOL_DIR/tests/generic/supported - - - - -``` - - -## 测试流程 -1. 除非指定具体工具名称,否则按顺序运行所有工具,遇到错误则停止运行。 -2. 先学习测试工具,再使用工具进行测试。 -3. 每个工具跑完之后,查看测试结果,如果测试结果不符合预期,可以根据日志信息进行排查,找出问题原因并尝试给出修复方案,等待确认。 -4. 如果有不确定的情况,可以咨询我,确保测试的正确性和有效性。 \ No newline at end of file diff --git a/scripts/dev-mds/run_all_test.sh b/scripts/dev-mds/run_all_test.sh index 243185a94..228d1868f 100755 --- a/scripts/dev-mds/run_all_test.sh +++ b/scripts/dev-mds/run_all_test.sh @@ -9,6 +9,7 @@ if [[ ! -d "$mydir" ]]; then mydir="$PWD"; fi DEFINE_string type 'all' 'test type' DEFINE_string mds_addr '' 'mds address' DEFINE_string mountpoint '' 'mount point' +DEFINE_integer round 1 'test round count' # parse the command-line @@ -30,7 +31,6 @@ fi BASE_DIR=$(dirname $(dirname $(cd $(dirname $0); pwd))) MOUNTPOINT=${FLAGS_mountpoint} -SUFFIX=$(date +%Y%m%d%H%M%S) function run_e2e_test() { echo "### [e2e] run test......" @@ -50,6 +50,15 @@ function run_e2e_test() { # run test command uv run pytest --mount-point=$TEST_ROOT_DIR > $E2E_LOG_DIR/e2e_test.log 2>&1 + # verify result + if grep -qE '[0-9]+ passed' $E2E_LOG_DIR/e2e_test.log && + ! grep -qE '[0-9]+ (failed|error)' $E2E_LOG_DIR/e2e_test.log; then + echo "### [e2e] result: PASS" + else + echo "### [e2e] result: FAIL" + FAILED=1 + fi + # uv run pytest quota --mount-point=$TEST_ROOT_DIR -m slow --mds-addr=${FLAGS_mds_addr} --fs-id=10000 --root-ino=1 echo "### [e2e] test done, log file: $E2E_LOG_DIR/e2e_test.log" @@ -73,6 +82,14 @@ function run_pjdtest_test() { # run test command sudo prove -rv --exec 'bash -x' ${PJD_DIR} > $PJD_LOG_DIR/pjd_test.log 2>&1 + # verify result + if grep -q '^Result: PASS' $PJD_LOG_DIR/pjd_test.log; then + echo "### [pjdtest] result: PASS" + else + echo "### [pjdtest] result: FAIL" + FAILED=1 + fi + echo "### [pjdtest] test done, log file: $PJD_LOG_DIR/pjd_test.log" } @@ -89,7 +106,15 @@ function run_fsx_test() { mkdir -p ${FSX_LOG_DIR} # run test command - fsx -l 1073741824 -o 1048576 -S 0 -p 10000 --duration=3600 --record-ops=$FSX_LOG_DIR/fsx.ops -P $FSX_LOG_DIR $FSX_TEST_FILE + fsx -l 1073741824 -o 1048576 -S 0 -p 10000 --duration=3600 --record-ops=$FSX_LOG_DIR/fsx.ops -P $FSX_LOG_DIR $FSX_TEST_FILE > $FSX_LOG_DIR/fsx.log 2>&1 + + # verify result + if grep -qE '^All [0-9]+ operations completed A-OK' $FSX_LOG_DIR/fsx.log; then + echo "### [fsx] result: PASS" + else + echo "### [fsx] result: FAIL" + FAILED=1 + fi echo "### [fsx] test done, log file: $FSX_LOG_DIR/fsx.ops" } @@ -173,18 +198,32 @@ function run_all_tests() { run_fsstress_test } -if [ "$FLAGS_type" = "all" ]; then - run_all_tests -elif [ "$FLAGS_type" = "e2e" ]; then - run_e2e_test -elif [ "$FLAGS_type" = "pjdtest" ]; then - run_pjdtest_test -elif [ "$FLAGS_type" = "fsx" ]; then - run_fsx_test -elif [ "$FLAGS_type" = "mdtest" ]; then - run_mdtest_test -elif [ "$FLAGS_type" = "fio" ]; then - run_fio_test -elif [ "$FLAGS_type" = "fsstress" ]; then - run_fsstress_test +FAILED=0 + +for ((i = 1; i <= ${FLAGS_round}; i++)); do + echo "### ===== round $i/${FLAGS_round} =====" + SUFFIX=$(date +%Y%m%d%H%M%S)_${i} + + if [ "$FLAGS_type" == "all" ]; then + run_all_tests + elif [ "$FLAGS_type" == "e2e" ]; then + run_e2e_test + elif [ "$FLAGS_type" == "pjdtest" ]; then + run_pjdtest_test + elif [ "$FLAGS_type" == "fsx" ]; then + run_fsx_test + elif [ "$FLAGS_type" == "mdtest" ]; then + run_mdtest_test + elif [ "$FLAGS_type" == "fio" ]; then + run_fio_test + elif [ "$FLAGS_type" == "fsstress" ]; then + run_fsstress_test + fi + + sleep 10 +done + +if [ ${FAILED} -ne 0 ]; then + echo "### some tests FAILED" + exit 1 fi \ No newline at end of file diff --git a/src/client/vfs/metasystem/mds/file_session.cc b/src/client/vfs/metasystem/mds/file_session.cc index aae94b666..0f90ff09e 100644 --- a/src/client/vfs/metasystem/mds/file_session.cc +++ b/src/client/vfs/metasystem/mds/file_session.cc @@ -236,7 +236,6 @@ FileSessionSPtr FileSessionMap::Put(Ino ino, uint64_t fh, } file_session = it->second; - // Keep handle registration atomic with the inode's last close. file_session->AddSession(fh, session_id, flags); }, ino); diff --git a/src/common/config_mapper.h b/src/common/config_mapper.h index cecd19062..f1f84d45e 100644 --- a/src/common/config_mapper.h +++ b/src/common/config_mapper.h @@ -35,6 +35,11 @@ static inline void FillBlockAccessOption( block_access_opt->s3_options.s3_info.sk = s3_info.sk(); block_access_opt->s3_options.s3_info.endpoint = s3_info.endpoint(); block_access_opt->s3_options.s3_info.bucket_name = s3_info.bucketname(); + // S3 backend config (log dir/prefix, crt client, timeouts, ...) comes + // from the --s3_* gflags; without this the AWS SDK falls back to an empty + // log prefix and writes /.log. + blockaccess::FillAwsSdkConfigFromGFlags( + &block_access_opt->s3_options.aws_sdk_config); } else if (fs_info.fs_type() == pb::mds::FsType::RADOS) { CHECK(fs_info.extra().has_rados_info()) << "ilegall storage_info, RADOS info not set"; diff --git a/src/mds/common/type.h b/src/mds/common/type.h index 5ffd6790a..eebcd30a4 100644 --- a/src/mds/common/type.h +++ b/src/mds/common/type.h @@ -15,8 +15,10 @@ #ifndef DINGOFS_MDS_COMMON_TYPE_H_ #define DINGOFS_MDS_COMMON_TYPE_H_ +#include #include +#include #include #include @@ -168,35 +170,183 @@ struct LocalFileInfo { std::string path; }; +struct AttrVersionVec; + struct AttrWithMutation { AttrEntry attr; absl::InlinedVector mutations; uint64_t BaseVersion() const { return attr.version(); } + uint64_t CompleteVersion() const { return attr.version() + TotalDeltaVersion(); } + uint64_t DeltaVersion(uint32_t index) const { + CHECK(index < mutations.size()) << "invalid mutation index(" << index << "), should be less than " + << mutations.size(); + return mutations[index].delta_version(); + } + // Mutations are per-slot absolute counters (see AttrVersionVec::PutIf). Callers + // may append them across retries, so dedup per slot instead of summing blindly: + // a duplicated slot must not inflate the version. uint64_t TotalDeltaVersion() const { - uint64_t total_delta_version = 0; + CHECK(mutations.size() <= kDirAttrMutationNum) << "too many mutations: " << mutations.size(); + + absl::InlinedVector deltas(kDirAttrMutationNum, 0); for (const auto& mutation : mutations) { - total_delta_version += mutation.delta_version(); + deltas[mutation.index()] = std::max(mutation.delta_version(), deltas[mutation.index()]); } + uint64_t total_delta_version = 0; + for (uint64_t delta : deltas) total_delta_version += delta; + return total_delta_version; } - uint64_t Version() const { return attr.version() + TotalDeltaVersion(); } - AttrEntry ToCompleteAttr() const { AttrEntry latest_attr = attr; for (const auto& mutation : mutations) { latest_attr.set_atime(std::max(latest_attr.atime(), mutation.atime())); latest_attr.set_mtime(std::max(latest_attr.mtime(), mutation.mtime())); latest_attr.set_ctime(std::max(latest_attr.ctime(), mutation.ctime())); - latest_attr.set_version(latest_attr.version() + mutation.delta_version()); } + latest_attr.set_version(attr.version() + TotalDeltaVersion()); + return latest_attr; } }; +// Represents a single attribute version, which can be either a base version or a delta version. +struct AttrVersion { + bool is_delta{false}; + uint32_t index{0}; + uint64_t version{0}; + AttrVersion(uint64_t version) : version(version) {} + AttrVersion(uint32_t index, uint64_t version) : is_delta(true), index(index), version(version) {} + + std::string ToString() const { + if (!is_delta) { + return fmt::format("{}", version); + } else { + return fmt::format("{}-{}", index, version); + } + } +}; + +// Represents a collection of attribute versions, including a base version and multiple delta versions. +struct AttrVersionVec { + // base version + uint64_t base_version{0}; + // delta versions + absl::InlinedVector delta_versions; + // sum of delta versions, used for quick check if there is mutation + uint64_t total_delta_version{0}; + + AttrVersionVec(uint64_t base_version) : base_version(base_version) { delta_versions.resize(kDirAttrMutationNum, 0); } + AttrVersionVec(const AttrWithMutation& attr_with_mutation) : base_version(attr_with_mutation.attr.version()) { + delta_versions.resize(kDirAttrMutationNum, 0); + total_delta_version = 0; + for (const auto& mutation : attr_with_mutation.mutations) { + PutIf(mutation); + } + } + + uint64_t BaseVersion() const { return base_version; } + uint64_t CompleteVersion() const { return base_version + total_delta_version; } + uint64_t DeltaVersion(uint32_t index) const { + CHECK(index < kDirAttrMutationNum) << fmt::format("out of range, {}/{}.", index, kDirAttrMutationNum); + + return delta_versions[index]; + } + + bool PutIf(const AttrMutationEntry& attr_mutation) { + uint32_t index = attr_mutation.index(); + CHECK(index < kDirAttrMutationNum) << fmt::format("out of range, {}/{}.", index, kDirAttrMutationNum); + + uint64_t& delta_version = delta_versions[index]; + if (attr_mutation.delta_version() <= delta_version) return false; + + total_delta_version += (attr_mutation.delta_version() - delta_version); + delta_version = attr_mutation.delta_version(); + + return true; + } + + bool PutIf(const AttrVersion& attr_version) { + if (!attr_version.is_delta) { + if (attr_version.version <= base_version) return false; + base_version = attr_version.version; + + } else { + CHECK(attr_version.index < kDirAttrMutationNum) + << fmt::format("out of range, {}/{}.", attr_version.index, kDirAttrMutationNum); + + uint64_t& delta_version = delta_versions[attr_version.index]; + if (attr_version.version <= delta_version) return false; + + total_delta_version += (attr_version.version - delta_version); + delta_version = attr_version.version; + } + + return true; + } + + bool PutIf(const AttrEntry& attr) { + if (attr.version() <= base_version) return false; + + base_version = attr.version(); + + return true; + } + + bool PutIf(const AttrWithMutation& attr_with_mutation) { + bool updated = false; + updated |= PutIf(attr_with_mutation.attr); + for (const auto& mutation : attr_with_mutation.mutations) { + updated |= PutIf(mutation); + } + + return updated; + } + + bool PutIf(const AttrVersionVec& version_vec) { + bool updated = false; + if (version_vec.base_version > base_version) { + updated = true; + base_version = version_vec.base_version; + } + + uint32_t size = std::min(delta_versions.size(), version_vec.delta_versions.size()); + for (uint32_t i = 0; i < size; ++i) { + if (version_vec.delta_versions[i] > delta_versions[i]) { + updated = true; + total_delta_version += (version_vec.delta_versions[i] - delta_versions[i]); + delta_versions[i] = version_vec.delta_versions[i]; + } + } + + return updated; + } + + // less than or equal + bool LessThanOrEqual(const AttrVersion& other) const { + if (!other.is_delta) { + return base_version <= other.version; + } else { + return DeltaVersion(other.index) <= other.version; + } + } + + // greater than or equal + bool GreaterThanOrEqual(const AttrVersion& other) const { + if (!other.is_delta) { + return base_version >= other.version; + } else { + return DeltaVersion(other.index) >= other.version; + } + } + + std::string ToString() const { return fmt::format("{} {}", base_version, total_delta_version); } +}; + // AttrOrMutation carries either a full parent inode attr (need_parent_key path) // or only a delta mutation against it (mutation path). The two are disambiguated // by attr.ino(): when attr.ino() == 0 the full attr was not loaded and only the @@ -207,6 +357,10 @@ struct AttrOrMutation { AttrMutationEntry mutation; bool IsMutation() const { return attr.ino() == 0; } + + AttrVersion ToAttrVersion() const { + return !IsMutation() ? AttrVersion(attr.version()) : AttrVersion(mutation.index(), mutation.delta_version()); + } }; enum class ReqType : uint8_t { diff --git a/src/mds/filesystem/filesystem.cc b/src/mds/filesystem/filesystem.cc index 75f979069..c5118286d 100644 --- a/src/mds/filesystem/filesystem.cc +++ b/src/mds/filesystem/filesystem.cc @@ -401,8 +401,7 @@ Status FileSystem::BatchCreate(Context& ctx, Ino parent, const std::vectorToAttr(); - for (auto& dentry : dentries) AddDentryToPartition(parent, dentry, last_parent_attr.version()); + for (auto& dentry : dentries) AddDentryToPartition(parent, dentry, parent_attr_or_mutation.ToAttrVersion()); // update quota quota_manager_.AsyncUpdateFsUsage(0, params.size(), reason); @@ -413,6 +412,7 @@ Status FileSystem::BatchCreate(Context& ctx, Ino parent, const std::vectorToAttr(); if (operation.GetBatchIndex() == 0 && IsParentHashPartition()) { NotifyBuddyRefreshInode(last_parent_attr.parents(), parent_attr_or_mutation, reason); } @@ -511,7 +511,7 @@ Status FileSystem::MkNod(Context& ctx, const MkNodParam& param, EntryWithPaOut& InsertInodeCache(attr, reason); // add dentry to partition AttrEntry last_parent_attr = parent_inode->ToAttr(); - AddDentryToPartition(parent, dentry, last_parent_attr.version()); + AddDentryToPartition(parent, dentry, parent_attr_or_mutation.ToAttrVersion()); // update quota quota_manager_.AsyncUpdateFsUsage(0, 1, reason); @@ -633,7 +633,7 @@ Status FileSystem::BatchMkNod(Context& ctx, const std::vector& param for (const auto& attr : attrs) InsertInodeCache(attr, reason); // add dentry to partition - for (const auto& dentry : dentries) AddDentryToPartition(parent, dentry, last_parent_attr.version()); + for (const auto& dentry : dentries) AddDentryToPartition(parent, dentry, parent_attr_or_mutation.ToAttrVersion()); // update quota quota_manager_.AsyncUpdateFsUsage(0, params.size(), reason); @@ -1111,7 +1111,8 @@ Status FileSystem::MkDir(Context& ctx, const MkDirParam& param, EntryWithPaOut& trace.RecordElapsedTime("post_inode"); // add dentry to partition - AddDentryToPartition(parent, dentry, last_parent_inode->Version()); + AttrVersion parent_attr_version(parent_attr.version()); + AddDentryToPartition(parent, dentry, parent_attr_version); trace.RecordElapsedTime("post_dentry"); @@ -1232,7 +1233,8 @@ Status FileSystem::BatchMkDir(Context& ctx, const std::vector& param for (auto& attr : attrs) InsertInodeCache(attr, reason); InodeSPtr last_parent_inode = UpsertInodeCache(parent_attr, reason); // add dentry to partition - for (auto& dentry : dentries) AddDentryToPartition(parent, dentry, last_parent_inode->Version()); + AttrVersion parent_attr_version(parent_attr.version()); + for (auto& dentry : dentries) AddDentryToPartition(parent, dentry, parent_attr_version); // update quota quota_manager_.AsyncUpdateFsUsage(0, params.size(), reason); @@ -1327,7 +1329,8 @@ Status FileSystem::RmDir(Context& ctx, Ino parent, const std::string& name, Entr std::string reason = fmt::format("rmdir.{}.{}.{}", request_id, parent, name); InodeSPtr last_parent_inode = UpsertInodeCache(parent_attr, reason); // delete dentry from partition - DeleteDentryFromPartition(parent, name, last_parent_inode->Version()); + AttrVersion parent_attr_version(parent_attr.version()); + DeleteDentryFromPartition(parent, name, parent_attr_version); if (enable_trash) { // Refresh child inode cache so immutability gates (CheckCreateInTrash etc.) // see parents=[trash_ino] on the hot path instead of a stale [orig_parent]. @@ -1521,8 +1524,7 @@ Status FileSystem::Link(Context& ctx, Ino ino, Ino new_parent, const std::string // update inode cache InsertInodeCache(attr, reason); // add dentry to partition - AttrEntry last_parent_attr = parent_inode->ToAttr(); - AddDentryToPartition(new_parent, dentry, last_parent_attr.version()); + AddDentryToPartition(new_parent, dentry, parent_attr_or_mutation.ToAttrVersion()); // update quota quota_manager_.AsyncUpdateDirUsage(new_parent, attr.length(), 1, reason); @@ -1531,6 +1533,7 @@ Status FileSystem::Link(Context& ctx, Ino ino, Ino new_parent, const std::string reason); // notify buddy + AttrEntry last_parent_attr = parent_inode->ToAttr(); if (operation.GetBatchIndex() == 0 && IsParentHashPartition()) { NotifyBuddyRefreshInode(last_parent_attr.parents(), parent_attr_or_mutation, reason); } @@ -1616,7 +1619,7 @@ Status FileSystem::UnLink(Context& ctx, Ino parent, const std::string& name, Ent UpsertInodeCache(attr, reason); // delete dentry from partition AttrEntry last_parent_attr = parent_inode->ToAttr(); - DeleteDentryFromPartition(parent, name, last_parent_attr.version()); + DeleteDentryFromPartition(parent, name, parent_attr_or_mutation.ToAttrVersion()); // update quota: // - plain unlink: debit fs-level + per-dir immediately. @@ -1746,7 +1749,7 @@ Status FileSystem::BatchUnLink(Context& ctx, Ino parent, const std::vectorToAttr(); - DeleteDentryFromPartition(parent, names, last_parent_attr.version()); + DeleteDentryFromPartition(parent, names, parent_attr_or_mutation.ToAttrVersion()); if (enable_trash) RecordTrashMoveOutcome(trash.bucket_ino); @@ -1885,8 +1888,7 @@ Status FileSystem::Symlink(Context& ctx, const std::string& symlink, Ino new_par // update inode cache InsertInodeCache(attr, reason); // add dentry to partition - AttrEntry last_parent_attr = parent_inode->ToAttr(); - AddDentryToPartition(new_parent, dentry, last_parent_attr.version()); + AddDentryToPartition(new_parent, dentry, parent_attr_or_mutation.ToAttrVersion()); // update quota quota_manager_.AsyncUpdateFsUsage(0, 1, reason); @@ -1896,6 +1898,7 @@ Status FileSystem::Symlink(Context& ctx, const std::string& symlink, Ino new_par AsyncUpdateDirStat(new_parent, 0, 1, 0, reason); // note buddy + AttrEntry last_parent_attr = parent_inode->ToAttr(); if (operation.GetBatchIndex() == 0 && IsParentHashPartition()) { NotifyBuddyRefreshInode(last_parent_attr.parents(), parent_attr_or_mutation, reason); } @@ -2009,7 +2012,10 @@ Status FileSystem::SetAttr(Context& ctx, Ino ino, const SetAttrParam& param, Ent entry_out.expand_file = (delta_bytes > 0) ? true : false; entry_out.chunks.swap(effected_chunks); - if (IsDir(ino)) RefreshPartitionDeltaVersion(ino, entry_out.attr.version()); + if (IsDir(ino)) { + AttrVersion attr_version(attr.version()); + RefreshPartitionVersion(ino, attr_version); + } trace.RecordElapsedTime("post_handle"); @@ -2118,7 +2124,10 @@ Status FileSystem::SetXAttr(Context& ctx, Ino ino, const Inode::XAttrMap& xattrs // set output entry_out.attr = last_inode->ToAttr(); - if (IsDir(ino)) RefreshPartitionDeltaVersion(ino, entry_out.attr.version()); + if (IsDir(ino)) { + AttrVersion attr_version(attr.version()); + RefreshPartitionVersion(ino, attr_version); + } trace.RecordElapsedTime("post_handle"); @@ -2171,7 +2180,10 @@ Status FileSystem::RemoveXAttr(Context& ctx, Ino ino, const std::string& name, E // set output entry_out.attr = last_inode->ToAttr(); - if (IsDir(ino)) RefreshPartitionDeltaVersion(ino, entry_out.attr.version()); + if (IsDir(ino)) { + AttrVersion attr_version(attr.version()); + RefreshPartitionVersion(ino, attr_version); + } trace.RecordElapsedTime("post_handle"); @@ -2281,14 +2293,16 @@ Status FileSystem::Rename(Context& ctx, const RenameParam& param, RenameResult& if (IsMonoPartition()) { // old parent dentry/inode - DeleteDentryFromPartition(old_parent, old_name, out.old_parent_inode.version()); + AttrVersion old_parent_attr_version(old_parent_attr_with_mutation.BaseVersion()); + DeleteDentryFromPartition(old_parent, old_name, old_parent_attr_version); UpsertInodeCache(old_parent_attr_with_mutation, reason); // new parent dentry/inode auto new_parent_node = UpsertInodeCache(new_parent_attr_with_mutation, reason); Dentry new_dentry(fs_id_, new_name, new_parent, old_dentry.ino(), old_dentry.type(), 0); - AddDentryToPartition(new_parent, new_dentry, new_parent_node->Version()); + AttrVersion new_parent_attr_version(new_parent_attr_with_mutation.BaseVersion()); + AddDentryToPartition(new_parent, new_dentry, new_parent_attr_version); // delete exist new partition if (is_exist_new_dentry) { @@ -2305,14 +2319,15 @@ Status FileSystem::Rename(Context& ctx, const RenameParam& param, RenameResult& } } else { - // clean old parent partition cache - NotifyBuddyCleanPartitionCache(old_parent, reason); - // refresh new parent inode and dentry cache auto new_parent_inode = UpsertInodeCache(new_parent_attr_with_mutation, reason); - AddDentryToPartition(new_parent, new_dentry, new_parent_inode->Version()); + AttrVersion new_parent_attr_version(new_parent_attr_with_mutation.BaseVersion()); + AddDentryToPartition(new_parent, new_dentry, new_parent_attr_version); if (is_same_parent) { - DeleteDentryFromPartition(new_parent, old_dentry.name(), new_parent_inode->Version()); + DeleteDentryFromPartition(new_parent, old_dentry.name(), new_parent_attr_version); + } else { + // clean old parent partition cache + NotifyBuddyCleanPartitionCache(old_parent, reason); } // refresh parent of parent inode cache. kTrashInodeId is virtual and has @@ -2472,7 +2487,8 @@ Status FileSystem::RestoreFromTrash(Context& ctx, Ino trash_parent, const std::s // Add restored dentry to partition cache. Dentry dentry(fs_id_, actual_dst_name, actual_dst_parent, result.file_ino, result.file_type, 0); - AddDentryToPartition(actual_dst_parent, dentry, dst_parent_attr.version()); + AttrVersion dst_parent_attr_version(dst_parent_attr.version()); + AddDentryToPartition(actual_dst_parent, dentry, dst_parent_attr_version); // Push the fresh dst-parent/file attrs to the other MDSes caching them, // mirroring MkNod/Rename. Parent-hash only: under mono GetMdsIdByIno reads @@ -3784,7 +3800,7 @@ bool FileSystem::CanServe(uint64_t self_mds_id) { return false; } -void FileSystem::AddDentryToPartition(Ino parent, const Dentry& dentry, uint64_t version) { +void FileSystem::AddDentryToPartition(Ino parent, const Dentry& dentry, const AttrVersion& version) { // Trash parents (.trash root + hour buckets) never enter partition_cache_; // see FetchPartition for the design rationale. if (IsTrashInode(parent)) return; @@ -3797,7 +3813,7 @@ void FileSystem::AddDentryToPartition(Ino parent, const Dentry& dentry, uint64_t } } -void FileSystem::DeleteDentryFromPartition(Ino parent, const std::string& name, uint64_t version) { +void FileSystem::DeleteDentryFromPartition(Ino parent, const std::string& name, const AttrVersion& version) { if (IsTrashInode(parent)) return; auto partition = GetPartitionFromCache(parent); if (partition != nullptr) { @@ -3807,7 +3823,8 @@ void FileSystem::DeleteDentryFromPartition(Ino parent, const std::string& name, } } -void FileSystem::DeleteDentryFromPartition(Ino parent, const std::vector& names, uint64_t version) { +void FileSystem::DeleteDentryFromPartition(Ino parent, const std::vector& names, + const AttrVersion& version) { if (IsTrashInode(parent)) return; auto partition = GetPartitionFromCache(parent); if (partition != nullptr) { @@ -3817,18 +3834,18 @@ void FileSystem::DeleteDentryFromPartition(Ino parent, const std::vectorRefreshDeltaVersion(version); + if (partition != nullptr) partition->RefreshVersion(version); } Status FileSystem::GetPartition(Context& ctx, Ino parent, PartitionPtr& out_partition) { auto status = GetPartition(ctx, ctx.GetInodeVersion(), parent, out_partition); if (status.ok()) { LOG_DEBUG << fmt::format("[fs.{}.{}.{}] get partition({}/{}) this({}).", fs_id_, out_partition->INo(), - ctx.RequestId(), out_partition->BaseVersion(), out_partition->DeltaVersion(), + ctx.RequestId(), out_partition->BaseVersion(), out_partition->CompleteVersion(), (void*)out_partition.get()); } @@ -3869,7 +3886,7 @@ Status FileSystem::GetPartition(Context& ctx, uint64_t version, Ino parent, Part return status; } - uint64_t cache_version = use_base_version ? partition->BaseVersion() : partition->DeltaVersion(); + uint64_t cache_version = use_base_version ? partition->BaseVersion() : partition->CompleteVersion(); if (version > cache_version) { std::string reason = fmt::format("out-of-date.{}.{}.[{},cache{},req{}]", method_name, request_id, use_base_version, cache_version, version); @@ -3880,7 +3897,7 @@ Status FileSystem::GetPartition(Context& ctx, uint64_t version, Ino parent, Part // singleflight may have piggybacked on a fetch started before this // request; re-validate and refetch (as a new leader) if still stale. - uint64_t got_version = use_base_version ? out_partition->BaseVersion() : out_partition->DeltaVersion(); + uint64_t got_version = use_base_version ? out_partition->BaseVersion() : out_partition->CompleteVersion(); if (version > got_version) { status = FetchPartition(ctx, parent, reason + ".retry", out_partition); if (!status.ok()) { @@ -3973,17 +3990,19 @@ Status FileSystem::DoFetchPartition(Context& ctx, Ino parent, const std::string& auto status = RunOperation(&operation); if (!status.ok()) return status; auto attr_with_mutation = std::move(operation.GetResult().attr_with_mutation); - auto attr = attr_with_mutation.ToCompleteAttr(); + // auto attr = attr_with_mutation.ToCompleteAttr(); - auto partition = ShardPartition::New(operation_processor_, attr); + auto partition = ShardPartition::New(operation_processor_, attr_with_mutation); out_partition = partition_cache_.PutIf(partition); UpsertInodeCache(attr_with_mutation, reason); - LOG_DEBUG << fmt::format( - "[fs.{}.{}.{}.{}][{}us] fetch partition, version({}) shard_boundaries({}) reason({}).", fs_id_, parent, - method_name, request_id, duration.ElapsedUs(), attr.version(), - ::dingofs::Helper::VectorToString(::dingofs::Helper::PbRepeatedToVector(attr.shard_boundaries())), reason); + LOG_DEBUG << fmt::format("[fs.{}.{}.{}.{}][{}us] fetch partition, version({}) shard_boundaries({}) reason({}).", + fs_id_, parent, method_name, request_id, duration.ElapsedUs(), + attr_with_mutation.CompleteVersion(), + ::dingofs::Helper::VectorToString( + ::dingofs::Helper::PbRepeatedToVector(attr_with_mutation.attr.shard_boundaries())), + reason); return Status::OK(); } @@ -4072,7 +4091,7 @@ Status FileSystem::GetInode(Context& ctx, uint64_t version, Ino ino, InodeSPtr& return GetInodeFromStore(ctx, ino, reason, true, out_inode); } - uint64_t cache_version = use_base_version ? inode->BaseVersion() : inode->Version(); + uint64_t cache_version = use_base_version ? inode->BaseVersion() : inode->CompleteVersion(); if (cache_version < version) { std::string reason = fmt::format("out-of-date.{}.{}.[{},cache{},req{}]", method_name, request_id, use_base_version, cache_version, version); diff --git a/src/mds/filesystem/filesystem.h b/src/mds/filesystem/filesystem.h index a3d9e2dc1..fe2e27a37 100644 --- a/src/mds/filesystem/filesystem.h +++ b/src/mds/filesystem/filesystem.h @@ -370,6 +370,9 @@ class FileSystem : public std::enable_shared_from_this { Status DescribePartitionShard(Ino ino, Json::Value& value); + // for unit test + void Test_DeletePartitionFromCache(Ino parent) { partition_cache_.Delete(parent); } + private: friend class DebugServiceImpl; friend class FsStatServiceImpl; @@ -383,11 +386,11 @@ class FileSystem : public std::enable_shared_from_this { Status GenFileIno(Ino& ino); bool CanServe(uint64_t self_mds_id); - void AddDentryToPartition(Ino parent, const Dentry& dentry, uint64_t version); - void DeleteDentryFromPartition(Ino parent, const std::string& name, uint64_t version); - void DeleteDentryFromPartition(Ino parent, const std::vector& names, uint64_t version); + void AddDentryToPartition(Ino parent, const Dentry& dentry, const AttrVersion& version); + void DeleteDentryFromPartition(Ino parent, const std::string& name, const AttrVersion& version); + void DeleteDentryFromPartition(Ino parent, const std::vector& names, const AttrVersion& version); // for setattr/setxattr/removexattr, which may update dir attr but not change dentry - void RefreshPartitionDeltaVersion(Ino parent, uint64_t version); + void RefreshPartitionVersion(Ino parent, const AttrVersion& version); // get partition Status GetPartition(Context& ctx, Ino parent, PartitionPtr& out_partition); diff --git a/src/mds/filesystem/inode.cc b/src/mds/filesystem/inode.cc index 356b850ca..4384eecc4 100644 --- a/src/mds/filesystem/inode.cc +++ b/src/mds/filesystem/inode.cc @@ -58,17 +58,12 @@ void Inode::Put(const AttrEntry& attr) { xattrs_.emplace(xattr.first, xattr.second); } - base_version_ = attr.version(); - last_refresh_time_s_.store(utils::Timestamp(), std::memory_order_relaxed); } void Inode::ApplyMutation(const AttrWithMutation& attr_with_mutation) { for (const auto& mutation : attr_with_mutation.mutations) { - if (mutation.delta_version() <= delta_versions_[mutation.index()]) continue; - - total_delta_version_ += (mutation.delta_version() - delta_versions_[mutation.index()]); - delta_versions_[mutation.index()] = mutation.delta_version(); + if (!version_vec_.PutIf(mutation)) continue; ctime_ = std::max(ctime_, mutation.ctime()); mtime_ = std::max(mtime_, mutation.mtime()); @@ -77,26 +72,24 @@ void Inode::ApplyMutation(const AttrWithMutation& attr_with_mutation) { } void Inode::PutIf(const AttrEntry& attr, const std::string& reason) { - LOG_DEBUG << fmt::format("[inode.{}.{}] update attr, version({}-{}) in_base_version({}) reason({}).", fs_id_, ino_, - base_version_, total_delta_version_, attr.version(), reason); + LOG_DEBUG << fmt::format("[inode.{}.{}] update attr, version({}) in_base_version({}) reason({}).", fs_id_, ino_, + version_vec_.ToString(), attr.version(), reason); utils::WriteLockGuard lk(lock_); - if (attr.version() <= base_version_) return; - - Put(attr); + if (version_vec_.PutIf(attr)) Put(attr); } void Inode::PutIf(const AttrWithMutation& attr_with_mutation, const std::string& reason) { const auto& attr = attr_with_mutation.attr; - LOG_DEBUG << fmt::format("[inode.{}.{}] update attr with mutation, version({}-{}) in_version({}-{}) reason({}).", - fs_id_, ino_, base_version_, total_delta_version_, attr_with_mutation.BaseVersion(), + LOG_DEBUG << fmt::format("[inode.{}.{}] update attr with mutation, version({}) in_version({}-{}) reason({}).", fs_id_, + ino_, version_vec_.ToString(), attr_with_mutation.BaseVersion(), attr_with_mutation.TotalDeltaVersion(), reason); utils::WriteLockGuard lk(lock_); - if (attr.version() > base_version_) Put(attr); + if (version_vec_.PutIf(attr)) Put(attr); ApplyMutation(attr_with_mutation); } @@ -105,22 +98,19 @@ AttrEntry Inode::PutByMutation(const AttrMutationEntry& mutation, const std::str CHECK(mutation.index() < kDirAttrMutationNum) << fmt::format("invalid mutation index({}), should be less than {}.", mutation.index(), kDirAttrMutationNum); - LOG_DEBUG << fmt::format("[inode.{}.{}] update attr by mutation, version({}-{}) in_delta_version({}_{}) reason({}).", - fs_id_, ino_, base_version_, total_delta_version_, mutation.index(), - mutation.delta_version(), reason); + LOG_DEBUG << fmt::format("[inode.{}.{}] update attr by mutation, version({}) in_delta_version({}_{}) reason({}).", + fs_id_, ino_, version_vec_.ToString(), mutation.index(), mutation.delta_version(), reason); utils::WriteLockGuard lk(lock_); - uint64_t& delta_version = delta_versions_[mutation.index()]; - if (mutation.delta_version() <= delta_version) return ToAttrNoLock(); + if (mutation.delta_version() <= version_vec_.DeltaVersion(mutation.index())) return ToAttrNoLock(); + + version_vec_.PutIf(mutation); ctime_ = std::max(ctime_, mutation.ctime()); mtime_ = std::max(mtime_, mutation.mtime()); atime_ = std::max(atime_, mutation.atime()); - total_delta_version_ += (mutation.delta_version() - delta_version); - delta_version = mutation.delta_version(); - last_refresh_time_s_.store(utils::Timestamp(), std::memory_order_relaxed); return ToAttrNoLock(); @@ -168,7 +158,7 @@ Inode::AttrEntry Inode::ToAttrNoLock() { (*attr.mutable_xattrs())[key] = value; } - attr.set_version(base_version_ + total_delta_version_); + attr.set_version(version_vec_.CompleteVersion()); return attr; } diff --git a/src/mds/filesystem/inode.h b/src/mds/filesystem/inode.h index 8e13b1bf5..c9b735031 100644 --- a/src/mds/filesystem/inode.h +++ b/src/mds/filesystem/inode.h @@ -61,13 +61,10 @@ class Inode { symlink_(attr.symlink()), rdev_(attr.rdev()), flags_(attr.flags()), - base_version_(attr.version()), + version_vec_(attr.version()), parents_(attr.parents().begin(), attr.parents().end()) { last_active_time_s_ = utils::Timestamp(); - total_delta_version_ = 0; - delta_versions_.resize(kDirAttrMutationNum, 0); - for (const auto& xattr : attr.xattrs()) { xattrs_.emplace(xattr.first, xattr.second); } @@ -88,18 +85,11 @@ class Inode { symlink_(attr_with_mutation.attr.symlink()), rdev_(attr_with_mutation.attr.rdev()), flags_(attr_with_mutation.attr.flags()), - base_version_(attr_with_mutation.attr.version()), + version_vec_(attr_with_mutation), parents_(attr_with_mutation.attr.parents().begin(), attr_with_mutation.attr.parents().end()) { last_active_time_s_ = utils::Timestamp(); last_refresh_time_s_ = utils::Timestamp(); - total_delta_version_ = 0; - delta_versions_.resize(kDirAttrMutationNum, 0); - for (const auto& mutation : attr_with_mutation.mutations) { - delta_versions_[mutation.index()] = mutation.delta_version(); - total_delta_version_ += mutation.delta_version(); - } - for (const auto& xattr : attr_with_mutation.attr.xattrs()) { xattrs_.emplace(xattr.first, xattr.second); } @@ -167,11 +157,18 @@ class Inode { } uint64_t BaseVersion() const { utils::ReadLockGuard lk(lock_); - return base_version_; + return version_vec_.BaseVersion(); } - uint64_t Version() const { + uint64_t CompleteVersion() const { utils::ReadLockGuard lk(lock_); - return base_version_ + total_delta_version_; + return version_vec_.CompleteVersion(); + } + // Copy of the full version vector (base + per-bucket deltas). Mirrors + // ShardPartition::VersionVec() so callers can compare both sides bucket by + // bucket. Returns by value to stay race-free under the read lock. + AttrVersionVec VersionVec() const { + utils::ReadLockGuard lk(lock_); + return version_vec_; } std::vector Parents() const { @@ -239,13 +236,7 @@ class Inode { uint64_t mtime_{0}; uint64_t atime_{0}; - // base version - uint64_t base_version_{0}; - - // delta versions - absl::InlinedVector delta_versions_; - // sum of delta versions, used for quick check if there is mutation - uint64_t total_delta_version_{0}; + AttrVersionVec version_vec_; std::atomic last_active_time_s_{0}; std::atomic last_refresh_time_s_{0}; diff --git a/src/mds/filesystem/partition.cc b/src/mds/filesystem/partition.cc index b9df36415..6d7d24ca1 100644 --- a/src/mds/filesystem/partition.cc +++ b/src/mds/filesystem/partition.cc @@ -140,15 +140,15 @@ std::pair DirShard::Split(const std::string& key, ui Range left_range{range_.start, key}; Range right_range{key, range_.end}; - DirShardSPtr left_shard = DirShard::New(left_id, left_range, version_, std::move(left_dentries)); - DirShardSPtr right_shard = DirShard::New(right_id, right_range, version_, std::move(right_dentries)); + DirShardSPtr left_shard = DirShard::New(left_id, left_range, version_vec_, std::move(left_dentries)); + DirShardSPtr right_shard = DirShard::New(right_id, right_range, version_vec_, std::move(right_dentries)); return {left_shard, right_shard}; } std::string DirShard::ToString() const { return fmt::format("id({}) range[{},{}) version({}) size({})", id_, ::dingofs::Helper::StringToHex(range_.start), - ::dingofs::Helper::StringToHex(range_.end), version_, Size()); + ::dingofs::Helper::StringToHex(range_.end), version_vec_.ToString(), Size()); } bool DirShard::Empty() const { @@ -175,7 +175,7 @@ void DirShard::Dump(Json::Value& value) const { value["id"] = id_; value["start"] = ::dingofs::Helper::StringToHex(range_.start); value["end"] = ::dingofs::Helper::StringToHex(range_.end); - value["version"] = version_; + value["version"] = version_vec_.ToString(); value["size"] = children_.size(); } @@ -194,13 +194,19 @@ void DirShard::Snapshot(size_t offset, size_t limit, std::vector& dentri uint64_t ShardPartition::BaseVersion() { utils::ReadLockGuard lk(lock_); - return base_version_; + return version_vec_.BaseVersion(); } -uint64_t ShardPartition::DeltaVersion() { +uint64_t ShardPartition::CompleteVersion() { utils::ReadLockGuard lk(lock_); - return delta_version_; + return version_vec_.CompleteVersion(); +} + +AttrVersionVec ShardPartition::VersionVec() { + utils::ReadLockGuard lk(lock_); + + return version_vec_; } Status ShardPartition::Get(const std::string& name, Dentry& out) { @@ -272,31 +278,31 @@ Status ShardPartition::Scan(const std::string& trace_id, const std::string& star return Status::OK(); } -void ShardPartition::Put(const Dentry& dentry, uint64_t version) { +void ShardPartition::Put(const Dentry& dentry, const AttrVersion& version) { { utils::WriteLockGuard lk(lock_); - delta_version_ = std::max(version, delta_version_); - AddDeltaOpNoLock({DentryOpType::ADD, version, dentry, 0}); + version_vec_.PutIf(version); + AddDeltaOpNoLock(DentryOp(DentryOpType::ADD, version, dentry, 0)); } auto shard = GetShard(dentry.Name()); if (shard) shard->Put(dentry); } -void ShardPartition::Delete(const std::string& name, uint64_t version) { +void ShardPartition::Delete(const std::string& name, const AttrVersion& version) { { utils::WriteLockGuard lk(lock_); - delta_version_ = std::max(version, delta_version_); - AddDeltaOpNoLock({DentryOpType::DELETE, version, Dentry(name), 0}); + version_vec_.PutIf(version); + AddDeltaOpNoLock(DentryOp(DentryOpType::DELETE, version, Dentry(name), 0)); } auto shard = GetShard(name); if (shard) shard->Delete(name); } -void ShardPartition::Delete(const std::vector& names, uint64_t version) { +void ShardPartition::Delete(const std::vector& names, const AttrVersion& version) { for (const auto& name : names) { auto shard = GetShard(name); if (shard) shard->Delete(name); @@ -305,17 +311,17 @@ void ShardPartition::Delete(const std::vector& names, uint64_t vers { utils::WriteLockGuard lk(lock_); - delta_version_ = std::max(version, delta_version_); + version_vec_.PutIf(version); for (const auto& name : names) { - AddDeltaOpNoLock({DentryOpType::DELETE, version, Dentry(name), 0}); + AddDeltaOpNoLock(DentryOp(DentryOpType::DELETE, version, Dentry(name), 0)); } } } -void ShardPartition::RefreshDeltaVersion(uint64_t version) { +void ShardPartition::RefreshVersion(const AttrVersion& version) { utils::WriteLockGuard lk(lock_); - delta_version_ = std::max(version, delta_version_); + version_vec_.PutIf(version); } bool ShardPartition::NeedCompact() { @@ -383,8 +389,7 @@ void ShardPartition::Dump(Json::Value& value, size_t dentry_offset, size_t dentr value["fs_id"] = fs_id_; value["ino"] = ino_; - value["base_version"] = base_version_; - value["delta_version"] = delta_version_; + value["version"] = version_vec_.ToString(); value["delta_dentry_ops_count"] = delta_dentry_ops_.size(); std::map shard_map_value; @@ -422,7 +427,7 @@ void ShardPartition::Dump(Json::Value& value, size_t dentry_offset, size_t dentr auto& shard_value = it->second; shard_value["id"] = shard->ID(); shard_value["size"] = shard->Size(); - shard_value["version"] = shard->Version(); + shard_value["version"] = shard->VersionString(); loaded_shards.push_back(shard); } @@ -432,7 +437,7 @@ void ShardPartition::Dump(Json::Value& value, size_t dentry_offset, size_t dentr } value["shards"] = shards_value; - std::sort(loaded_shards.begin(), loaded_shards.end(), + std::sort(loaded_shards.begin(), loaded_shards.end(), // NOLINT [](const DirShardSPtr& lhs, const DirShardSPtr& rhs) { return lhs->Start() < rhs->Start(); }); size_t remaining_offset = dentry_offset; size_t total_dentries = 0; @@ -463,7 +468,7 @@ void ShardPartition::Dump(Json::Value& value, size_t dentry_offset, size_t dentr for (size_t i = 0; it != delta_dentry_ops_.end() && i < delta_limit; ++it, ++i) { Json::Value op_value(Json::objectValue); op_value["op"] = it->op_type == DentryOpType::ADD ? "ADD" : "DELETE"; - op_value["version"] = it->version; + op_value["version"] = it->version.ToString(); op_value["time_s"] = it->time_s; DentryToJson(it->dentry, op_value["dentry"]); delta_value.append(op_value); @@ -477,27 +482,11 @@ void ShardPartition::AddDeltaOpNoLock(DentryOp&& op) { op.time_s = utils::Timestamp(); delta_dentry_ops_.push_back(std::move(op)); - - // keep delta_dentry_ops_ ordered by version, the new op is usually with greater version, so we compare with the last - // op first to avoid unnecessary sort - if (delta_dentry_ops_.size() > 1) { - auto it = delta_dentry_ops_.end(); - --it; // last element - auto prev = it; - --prev; - while (true) { - if (it->version >= prev->version) break; - std::iter_swap(it, prev); - if (prev == delta_dentry_ops_.begin()) break; - it = prev; - --prev; - } - } } void ShardPartition::ApplyDeltaOpNoLock(DirShardSPtr shard) { for (auto& op : delta_dentry_ops_) { - if (op.version <= shard->Version()) continue; + if (shard->VersionVec().GreaterThanOrEqual(op.version)) continue; // check shard range contains dentry name, if not, skip this op and let it be applied to next shard after split if (!shard->Contains(op.dentry.Name())) continue; @@ -617,9 +606,9 @@ Status ShardPartition::DoFetchDirShard(const Range& range, const std::string& re auto& result = operation.GetResult(); - uint64_t version = result.attr_with_mutation.Version(); + AttrVersionVec version_vec(result.attr_with_mutation); - out_shard = DirShard::New(NextShardID(), range, version, std::move(dentries)); + out_shard = DirShard::New(NextShardID(), range, version_vec, std::move(dentries)); PutShard(out_shard); @@ -629,25 +618,24 @@ Status ShardPartition::DoFetchDirShard(const Range& range, const std::string& re return Status::OK(); } -bool ShardPartition::Refresh(uint64_t new_version) { - LOG_DEBUG << fmt::format("[partition.{}.{}] refresh partition, version({}->{}) delta_version({}).", fs_id_, ino_, - base_version_, new_version, delta_version_); +bool ShardPartition::Refresh(const AttrVersionVec& version_vec) { + LOG_DEBUG << fmt::format("[partition.{}.{}] refresh partition, version({}->{}).", fs_id_, ino_, + version_vec_.ToString(), version_vec.ToString()); utils::WriteLockGuard lk(lock_); - if (new_version <= base_version_) return false; + if (!version_vec_.PutIf(version_vec)) return false; - base_version_ = new_version; - delta_version_ = std::max(delta_version_, base_version_); shard_map_.clear(); for (auto it = delta_dentry_ops_.begin(); it != delta_dentry_ops_.end();) { - if (it->version <= base_version_) { + if (version_vec_.GreaterThanOrEqual(it->version)) { it = delta_dentry_ops_.erase(it); continue; } - delta_version_ = std::max(delta_version_, it->version); + version_vec_.PutIf(it->version); + ++it; } @@ -794,7 +782,7 @@ PartitionPtr PartitionCache::PutIf(const PartitionPtr& partition) { total_count_ << 1; } else { - it->second->Refresh(partition->BaseVersion()); + it->second->Refresh(partition->VersionVec()); new_partition = it->second; } }, diff --git a/src/mds/filesystem/partition.h b/src/mds/filesystem/partition.h index af8c9ab61..1a2e7b8b2 100644 --- a/src/mds/filesystem/partition.h +++ b/src/mds/filesystem/partition.h @@ -40,8 +40,8 @@ using DirShardSPtr = std::shared_ptr; class DirShard { public: - DirShard(uint64_t id, const Range& range, uint64_t version, const std::vector& dentries) - : id_(id), range_{range}, version_(version) { + DirShard(uint64_t id, const Range& range, const AttrVersionVec& version_vec, const std::vector& dentries) + : id_(id), range_{range}, version_vec_(version_vec) { // ingest dentries to map for (const auto& dentry : dentries) { CHECK(Contains(dentry.Name())) << fmt::format("dentry name({}) out of shard range{}.", dentry.Name(), @@ -51,21 +51,23 @@ class DirShard { last_active_time_s_ = utils::Timestamp(); last_refresh_time_s_ = utils::Timestamp(); } - DirShard(uint64_t id, const Range& range, uint64_t version, absl::btree_map&& dentries) - : id_(id), range_{range}, version_(version) { + DirShard(uint64_t id, const Range& range, const AttrVersionVec& version_vec, + absl::btree_map&& dentries) + : id_(id), range_{range}, version_vec_(version_vec) { // ingest dentries to map children_ = std::move(dentries); last_active_time_s_ = utils::Timestamp(); last_refresh_time_s_ = utils::Timestamp(); } - static DirShardSPtr New(uint64_t id, const Range& range, uint64_t version, const std::vector& dentries) { - return std::make_shared(id, range, version, dentries); + static DirShardSPtr New(uint64_t id, const Range& range, const AttrVersionVec& version_vec, + const std::vector& dentries) { + return std::make_shared(id, range, version_vec, dentries); } - static DirShardSPtr New(uint64_t id, const Range& range, uint64_t version, + static DirShardSPtr New(uint64_t id, const Range& range, const AttrVersionVec& version_vec, absl::btree_map&& dentries) { - return std::make_shared(id, range, version, std::move(dentries)); + return std::make_shared(id, range, version_vec, std::move(dentries)); } uint64_t ID() const { return id_; } @@ -96,7 +98,8 @@ class DirShard { void UpdateLastRefreshTime() { last_refresh_time_s_.store(utils::Timestamp(), std::memory_order_relaxed); } uint64_t LastRefreshTimeS() { return last_refresh_time_s_.load(std::memory_order_relaxed); } - uint64_t Version() const { return version_; } + const AttrVersionVec& VersionVec() { return version_vec_; } + std::string VersionString() const { return version_vec_.ToString(); } std::pair Split(const std::string& key, uint64_t left_id, uint64_t right_id); @@ -112,7 +115,8 @@ class DirShard { private: const uint64_t id_; const Range range_; // [start, end) - const uint64_t version_; + + const AttrVersionVec version_vec_; mutable utils::RWLock lock_; absl::btree_map children_; @@ -126,16 +130,24 @@ using PartitionPtr = std::shared_ptr; class ShardPartition { public: - ShardPartition(OperationProcessorSPtr operation_processor, const AttrEntry& attr) + ShardPartition(OperationProcessorSPtr operation_processor, const AttrEntry attr) : fs_id_(attr.fs_id()), ino_(attr.ino()), - base_version_(attr.version()), - delta_version_(base_version_), + version_vec_(attr.version()), operation_processor_(operation_processor) { for (const auto& boundary : attr.shard_boundaries()) { shard_boundaries_.push_back(boundary); } } + ShardPartition(OperationProcessorSPtr operation_processor, const AttrWithMutation& attr_with_mutation) + : fs_id_(attr_with_mutation.attr.fs_id()), + ino_(attr_with_mutation.attr.ino()), + version_vec_(attr_with_mutation), + operation_processor_(operation_processor) { + for (const auto& boundary : attr_with_mutation.attr.shard_boundaries()) { + shard_boundaries_.push_back(boundary); + } + } ~ShardPartition() = default; @@ -143,11 +155,16 @@ class ShardPartition { return std::make_shared(operation_processor, attr); } + static PartitionPtr New(OperationProcessorSPtr operation_processor, const AttrWithMutation& attr_with_mutation) { + return std::make_shared(operation_processor, attr_with_mutation); + } + uint32_t FsId() const { return fs_id_; } Ino INo() const { return ino_; } uint64_t BaseVersion(); - uint64_t DeltaVersion(); + uint64_t CompleteVersion(); + AttrVersionVec VersionVec(); Status Get(const std::string& name, Dentry& out); std::vector GetAll(); @@ -155,15 +172,20 @@ class ShardPartition { Status Scan(const std::string& trace_id, const std::string& start_name, uint32_t limit, bool is_only_dir, std::vector& dentries); - void Put(const Dentry& dentry, uint64_t version); - void Delete(const std::string& name, uint64_t version); - void Delete(const std::vector& names, uint64_t version); + void Put(const Dentry& dentry, const AttrVersion& version); + void Delete(const std::string& name, const AttrVersion& version); + void Delete(const std::vector& names, const AttrVersion& version); // refresh partition with latest version, for SetAttr/SetXAttr/RemoveXAttr - void RefreshDeltaVersion(uint64_t version); + void RefreshVersion(const AttrVersion& version); // delta dentry op too many may cause performance issue, need compact to reduce the op count. bool NeedCompact(); + void TEST_DeleteDirShard() { + utils::WriteLockGuard lk(lock_); + shard_map_.clear(); + } + bool Empty() const; size_t Size() const; size_t ShardSize() const; @@ -180,9 +202,11 @@ class ShardPartition { struct DentryOp { DentryOpType op_type; - uint64_t version; + AttrVersion version; Dentry dentry; uint64_t time_s; + DentryOp(DentryOpType op_type, const AttrVersion& version, const Dentry& dentry, uint64_t time_s) + : op_type(op_type), version(version), dentry(dentry), time_s(time_s) {} }; void AddDeltaOpNoLock(DentryOp&& op); @@ -207,7 +231,7 @@ class ShardPartition { Status DoFetchDirShard(const Range& range, const std::string& reason, DirShardSPtr& out_shard); // refresh partition with latest inode - bool Refresh(uint64_t new_version); + bool Refresh(const AttrVersionVec& version_vec); // split dir shard Status DoSplitDirShard(const Range& range); @@ -220,8 +244,8 @@ class ShardPartition { mutable utils::RWLock lock_; - uint64_t base_version_{0}; - uint64_t delta_version_{0}; + AttrVersionVec version_vec_; + std::list delta_dentry_ops_; std::vector shard_boundaries_; diff --git a/src/mds/filesystem/store_operation.cc b/src/mds/filesystem/store_operation.cc index b35c11cc8..87a1ece74 100644 --- a/src/mds/filesystem/store_operation.cc +++ b/src/mds/filesystem/store_operation.cc @@ -2582,6 +2582,11 @@ Status CleanTrashBucketOperation::Run(TxnUPtr& txn) { } Status RenameOperation::Run(TxnUPtr& txn) { + // RunAlone retries Run on conflict; result_ must be rebuilt from scratch so + // mutations are not appended twice (which would inflate ToCompleteAttr()). + result_.old_parent_attr_with_mutation.mutations.clear(); + result_.new_parent_attr_with_mutation.mutations.clear(); + uint64_t time_ns = GetTime(); LOG_DEBUG << fmt::format("[operation.{}] rename old_parent({}), old_name({}), new_parent_ino({}), new_name({}).", @@ -3854,6 +3859,10 @@ Status ScanTrashDentryOperation::Run(TxnUPtr& txn) { } Status ScanDirShardOperation::Run(TxnUPtr& txn) { + // RunAlone retries Run on conflict; clear appended mutations so a retry does + // not accumulate duplicate slots. + result_.attr_with_mutation.mutations.clear(); + const Range complete_range = MetaCodec::GetDentryRange(fs_id_, ino_, false); Range range; @@ -4034,6 +4043,10 @@ Status GetInodeAttrOperation::Run(TxnUPtr& txn) { CHECK(fs_id_ > 0) << "fs_id is 0"; CHECK(ino_ > 0) << "ino is 0"; + // RunAlone retries Run on conflict; clear appended mutations so a retry does + // not accumulate duplicate slots. + result_.attr_with_mutation.mutations.clear(); + Status status; if (IsDir(ino_)) { Range range; diff --git a/src/mds/filesystem/store_operation.h b/src/mds/filesystem/store_operation.h index 779c5a827..5911be49a 100644 --- a/src/mds/filesystem/store_operation.h +++ b/src/mds/filesystem/store_operation.h @@ -597,6 +597,7 @@ class MkDirOperation : public Operation { : Operation(trace), dentry_(dentry), attr_(attr) {}; ~MkDirOperation() override = default; + // no use mutation, because mkdir need update parent nlink struct Result { AttrEntry parent_attr; }; @@ -627,6 +628,7 @@ class BatchMkDirOperation : public Operation { : Operation(trace), dentries_(dentries), attrs_(attrs) {}; ~BatchMkDirOperation() override = default; + // no use mutation, because mkdir need update parent nlink struct Result { AttrEntry parent_attr; }; @@ -844,6 +846,7 @@ class UpdateAttrOperation : public Operation { : Operation(trace), ino_(ino), to_set_(to_set), attr_(attr), extra_param_(extra_param) {}; ~UpdateAttrOperation() override = default; + // not use mutation struct Result { AttrEntry attr; int64_t delta_bytes{0}; @@ -881,6 +884,7 @@ class UpdateXAttrOperation : public Operation { : Operation(trace), fs_id_(fs_id), ino_(ino), xattrs_(xattrs) {}; ~UpdateXAttrOperation() override = default; + // not use mutation struct Result { AttrEntry attr; }; @@ -910,6 +914,7 @@ class RemoveXAttrOperation : public Operation { : Operation(trace), fs_id_(fs_id), ino_(ino), name_(name) {}; ~RemoveXAttrOperation() override = default; + // not use mutation struct Result { AttrEntry attr; }; diff --git a/src/mds/service/fsstat_service.cc b/src/mds/service/fsstat_service.cc index 20451365f..2335dad9a 100644 --- a/src/mds/service/fsstat_service.cc +++ b/src/mds/service/fsstat_service.cc @@ -681,7 +681,7 @@ static void RenderPartitionCacheListPage(uint32_t fs_id, size_t page, std::vecto os << fmt::format( R"({})" R"({}{}{}{})", - fs_id, partition->INo(), partition->INo(), partition->BaseVersion(), partition->DeltaVersion(), + fs_id, partition->INo(), partition->INo(), partition->BaseVersion(), partition->CompleteVersion(), partition->ShardSize(), partition->Size()); } os << ""; @@ -708,7 +708,7 @@ static void RenderInodeCacheListPage(uint32_t fs_id, size_t page, std::vector{})" R"({}{}{})", - fs_id, inode->Ino(), inode->Ino(), pb::mds::FileType_Name(inode->Type()), inode->Version(), + fs_id, inode->Ino(), inode->Ino(), pb::mds::FileType_Name(inode->Type()), inode->CompleteVersion(), inode->IsFresh() ? "yes" : "no"); } os << ""; @@ -742,8 +742,7 @@ static void RenderChunkCacheListPage(uint32_t fs_id, size_t page, std::vectorPartition ino {}: base version {}, delta version {}", ino, - value["base_version"].asUInt64(), value["delta_version"].asUInt64()); + os << fmt::format("

Partition ino {}: version {}

", ino, value["version"].asString()); os << fmt::format(R"(

Shards [{}]

)" R"()", @@ -797,7 +796,7 @@ static void RenderInodeCachePage(uint32_t fs_id, Ino ino, const InodeSPtr& inode std::string attr_json; ::dingofs::Helper::ProtoToJson(attr, attr_json); os << fmt::format("

Inode {}: base version {}, current version {}, fresh {}

", ino, inode->BaseVersion(), - inode->Version(), inode->IsFresh() ? "yes" : "no"); + inode->CompleteVersion(), inode->IsFresh() ? "yes" : "no"); os << fmt::format("

Last active: {}
Last refresh: {}

", utils::FormatMsTime(inode->LastActiveTimeS() * 1000), utils::FormatMsTime(inode->LastRefreshTimeS() * 1000)); os << "

Attributes

" << HtmlEscape(attr_json) << "
"; diff --git a/src/mds/storage/dingodb_storage.cc b/src/mds/storage/dingodb_storage.cc index 8bf3f4f67..43189cec4 100644 --- a/src/mds/storage/dingodb_storage.cc +++ b/src/mds/storage/dingodb_storage.cc @@ -146,6 +146,8 @@ DingodbStorage::SdkTxnUPtr DingodbStorage::NewSdkTxn(Txn::IsolationLevel isolati LOG(ERROR) << fmt::format("[storage] new transaction fail, retry({}) error({}).", retry, status.ToString()); } while (++retry <= FLAGS_mds_txn_max_retry_times); + if (txn == nullptr) return nullptr; + return DingodbStorage::SdkTxnUPtr(txn); } @@ -374,7 +376,10 @@ Status DingodbStorage::Delete(const std::vector& keys) { } TxnUPtr DingodbStorage::NewTxn(Txn::IsolationLevel isolation_level) { - return std::make_unique(NewSdkTxn(isolation_level)); + auto sdk_txn = NewSdkTxn(isolation_level); + if (sdk_txn == nullptr) return nullptr; + + return std::make_unique(std::move(sdk_txn)); } DingodbTxn::~DingodbTxn() { diff --git a/src/mds/storage/dummy_storage.cc b/src/mds/storage/dummy_storage.cc index 3f114bf33..064a51221 100644 --- a/src/mds/storage/dummy_storage.cc +++ b/src/mds/storage/dummy_storage.cc @@ -22,11 +22,15 @@ #include #include +#include "common/helper.h" #include "dingofs/error.pb.h" +#include "fmt/format.h" namespace dingofs { namespace mds { +static void RandomSleep() { bthread_usleep(dingofs::Helper::GenerateRealRandomInteger(500, 3000)); } + bool DummyStorage::Init(const std::string&) { return true; } Status DummyStorage::CreateTable(const std::string& name, const TableOption& option, int64_t& table_id) { @@ -201,7 +205,8 @@ TxnUPtr DummyStorage::NewTxn(Txn::IsolationLevel isolation_level) { } Status DummyStorage::ApplyTxn(const std::map& writes, - const std::set& if_absent_keys) { + const std::set& if_absent_keys, + const std::map& read_versions) { utils::WriteLockGuard lock(lock_); // Re-verify if-absent keys atomically with the apply. @@ -216,12 +221,81 @@ Status DummyStorage::ApplyTxn(const std::map& writes, } } + // Optimistic concurrency: reject a stale read-modify-write. Only keys the txn + // read *and* writes are checked; pure blind writes keep last-write-wins. + for (const auto& [key, kv] : writes) { + auto rit = read_versions.find(key); + if (rit == read_versions.end()) continue; + + uint64_t current_version = VersionNoLock(key); + if (current_version != rit->second) { + return Status( + pb::error::ESTORE_MAYBE_RETRY, + fmt::format("write conflict on key, read_version({}) current_version({})", rit->second, current_version)); + } + } + + if (writes.empty()) return Status::OK(); + + uint64_t new_version = ++commit_version_; for (const auto& [key, kv] : writes) { if (kv.opt_type == KeyValue::OpType::kPut) { data_[key] = kv.value; } else if (kv.opt_type == KeyValue::OpType::kDelete) { data_.erase(key); } + key_versions_[key] = new_version; + } + + return Status::OK(); +} + +uint64_t DummyStorage::VersionNoLock(const std::string& key) const { + // A key that is absent from data_ reads as version 0 even if key_versions_ + // still holds the version of a since-deleted write; otherwise a + // read-absent-then-write (e.g. recreate after delete) would conflict forever. + if (data_.find(key) == data_.end()) return 0; + + auto it = key_versions_.find(key); + return it == key_versions_.end() ? 0 : it->second; +} + +Status DummyStorage::GetWithVersion(const std::string& key, std::string& value, uint64_t& version) { + utils::ReadLockGuard lock(lock_); + + auto it = data_.find(key); + if (it == data_.end()) { + version = 0; + return Status(pb::error::ENOT_FOUND, "key not found"); + } + + value = it->second; + version = VersionNoLock(key); + return Status::OK(); +} + +Status DummyStorage::BatchGetWithVersion(const std::vector& keys, std::vector& kvs, + std::map& versions) { + utils::ReadLockGuard lock(lock_); + + for (const auto& key : keys) { + auto it = data_.find(key); + if (it == data_.end()) continue; + + kvs.push_back(KeyValue{KeyValue::OpType::kPut, key, it->second}); + versions[key] = VersionNoLock(key); + } + + return Status::OK(); +} + +Status DummyStorage::ScanWithVersion(const Range& range, std::vector& kvs, + std::map& versions) { + utils::ReadLockGuard lock(lock_); + + for (auto it = data_.lower_bound(range.start); it != data_.end() && it->first < range.end; ++it) { + kvs.push_back(KeyValue{KeyValue::OpType::kPut, it->first, it->second}); + versions[it->first] = VersionNoLock(it->first); } return Status::OK(); @@ -236,6 +310,8 @@ DummyTxn::DummyTxn(DummyStorage* storage, Txn::IsolationLevel isolation_level) int64_t DummyTxn::ID() const { return txn_id_; } Status DummyTxn::Put(const std::string& key, const std::string& value) { + RandomSleep(); + if (committed_) return Status(pb::error::EINTERNAL, "txn already committed"); stage_writes_[key] = KeyValue{KeyValue::OpType::kPut, key, value}; @@ -245,6 +321,8 @@ Status DummyTxn::Put(const std::string& key, const std::string& value) { } Status DummyTxn::PutIfAbsent(const std::string& key, const std::string& value) { + RandomSleep(); + if (committed_) return Status(pb::error::EINTERNAL, "txn already committed"); // A prior staged Put in this txn means the key is already present from the @@ -270,6 +348,8 @@ Status DummyTxn::PutIfAbsent(const std::string& key, const std::string& value) { } Status DummyTxn::Delete(const std::string& key) { + RandomSleep(); + if (committed_) return Status(pb::error::EINTERNAL, "txn already committed"); stage_writes_[key] = KeyValue{KeyValue::OpType::kDelete, key, ""}; @@ -278,6 +358,8 @@ Status DummyTxn::Delete(const std::string& key) { } Status DummyTxn::Get(const std::string& key, std::string& value) { + RandomSleep(); + auto it = stage_writes_.find(key); if (it != stage_writes_.end()) { if (it->second.opt_type == KeyValue::OpType::kDelete) { @@ -287,10 +369,17 @@ Status DummyTxn::Get(const std::string& key, std::string& value) { return Status::OK(); } - return storage_->Get(key, value); + uint64_t version = 0; + auto status = storage_->GetWithVersion(key, value, version); + if (status.ok() || status.error_code() == pb::error::ENOT_FOUND) { + read_versions_[key] = version; + } + return status; } Status DummyTxn::BatchGet(const std::vector& keys, std::vector& kvs) { + RandomSleep(); + std::vector rest_keys; rest_keys.reserve(keys.size()); @@ -309,13 +398,29 @@ Status DummyTxn::BatchGet(const std::vector& keys, std::vector rest_kvs; - auto status = storage_->BatchGet(rest_keys, rest_kvs); + std::map versions; + auto status = storage_->BatchGetWithVersion(rest_keys, rest_kvs, versions); if (!status.ok()) return status; + // Absent keys are version 0 so a later create is detected as a conflict. + for (const auto& key : rest_keys) { + auto vit = versions.find(key); + read_versions_[key] = (vit == versions.end()) ? 0 : vit->second; + } + kvs.insert(kvs.end(), rest_kvs.begin(), rest_kvs.end()); return Status::OK(); } +void DummyTxn::RecordReadVersions(const std::map& versions) { + for (const auto& [key, version] : versions) { + // Staged keys are this txn's own view, not storage state; skip them. + if (stage_writes_.find(key) != stage_writes_.end()) continue; + // Last read wins: the value this txn may build its write on. + read_versions_[key] = version; + } +} + // Merges a sorted storage scan result with the txn's staged writes for the // given range. Iterates in ascending key order, applying staged // puts/deletes to mask or override storage values. @@ -360,20 +465,30 @@ static void MergeScanWithStage(const Range& range, const std::map& kvs) { + RandomSleep(); + std::vector storage_kvs; - auto status = storage_->Scan(range, storage_kvs); + std::map versions; + auto status = storage_->ScanWithVersion(range, storage_kvs, versions); if (!status.ok()) return status; + RecordReadVersions(versions); + if (limit == 0) limit = UINT64_MAX; MergeScanWithStage(range, stage_writes_, std::move(storage_kvs), limit, kvs); return Status::OK(); } Status DummyTxn::Scan(const Range& range, ScanHandlerType handler) { + RandomSleep(); + std::vector storage_kvs; - auto status = storage_->Scan(range, storage_kvs); + std::map versions; + auto status = storage_->ScanWithVersion(range, storage_kvs, versions); if (!status.ok()) return status; + RecordReadVersions(versions); + std::vector merged; MergeScanWithStage(range, stage_writes_, std::move(storage_kvs), UINT64_MAX, merged); @@ -384,10 +499,15 @@ Status DummyTxn::Scan(const Range& range, ScanHandlerType handler) { } Status DummyTxn::Scan(const Range& range, std::function handler) { + RandomSleep(); + std::vector storage_kvs; - auto status = storage_->Scan(range, storage_kvs); + std::map versions; + auto status = storage_->ScanWithVersion(range, storage_kvs, versions); if (!status.ok()) return status; + RecordReadVersions(versions); + std::vector merged; MergeScanWithStage(range, stage_writes_, std::move(storage_kvs), UINT64_MAX, merged); @@ -401,9 +521,10 @@ Status DummyTxn::Commit() { if (committed_) return Status(pb::error::EINTERNAL, "txn already committed"); committed_ = true; - auto status = storage_->ApplyTxn(stage_writes_, if_absent_keys_); + auto status = storage_->ApplyTxn(stage_writes_, if_absent_keys_, read_versions_); stage_writes_.clear(); if_absent_keys_.clear(); + read_versions_.clear(); return status; } diff --git a/src/mds/storage/dummy_storage.h b/src/mds/storage/dummy_storage.h index 97326b0a6..1a07caff0 100644 --- a/src/mds/storage/dummy_storage.h +++ b/src/mds/storage/dummy_storage.h @@ -23,6 +23,7 @@ #include #include "absl/container/btree_map.h" +#include "absl/container/flat_hash_map.h" #include "mds/storage/storage.h" #include "utils/concurrent/concurrent.h" @@ -60,10 +61,25 @@ class DummyStorage : public KVStorage { private: friend class DummyTxn; - // Atomically verifies that none of `if_absent_keys` already exist, then - // applies `writes` (last-write-wins per key). Used by DummyTxn::Commit to - // make PutIfAbsent semantics race-free against concurrent committers. - Status ApplyTxn(const std::map& writes, const std::set& if_absent_keys); + // Atomically applies `writes` under the storage lock. Before applying, + // re-verifies `if_absent_keys` and, for every key the txn both read and is + // about to write, checks that it did not change since the read. A changed key + // yields ESTORE_MAYBE_RETRY, mirroring the optimistic concurrency of the + // production backends: without it a stale read-modify-write silently loses an + // update (last-write-wins). + Status ApplyTxn(const std::map& writes, const std::set& if_absent_keys, + const std::map& read_versions); + + // Versioned reads: the version is captured atomically with the value so a + // later commit can tell whether the key moved under the transaction. Absent + // keys are reported with version 0 (and ENOT_FOUND for the single-key read). + Status GetWithVersion(const std::string& key, std::string& value, uint64_t& version); + Status BatchGetWithVersion(const std::vector& keys, std::vector& kvs, + std::map& versions); + Status ScanWithVersion(const Range& range, std::vector& kvs, std::map& versions); + + // Caller must hold lock_. Returns 0 for a key that is absent or never written. + uint64_t VersionNoLock(const std::string& key) const; struct Table { std::string name; @@ -77,6 +93,11 @@ class DummyStorage : public KVStorage { std::map tables_; absl::btree_map data_; + + // Last commit version that wrote each key; the basis for write-conflict + // detection. Entries are never evicted (test-scope storage). + absl::flat_hash_map key_versions_; + uint64_t commit_version_{0}; }; class DummyTxn : public Txn { @@ -114,7 +135,13 @@ class DummyTxn : public Txn { // storage must be re-verified atomically at Commit time. std::set if_absent_keys_; + // Version observed the last time each key was read from storage. Commit + // rejects the txn if one of these keys changed before we wrote it. + std::map read_versions_; + bool committed_{false}; + + void RecordReadVersions(const std::map& versions); }; } // namespace mds diff --git a/src/tools/mds-cli/br.cc b/src/tools/mds-cli/br.cc index 542708827..8a0549e59 100644 --- a/src/tools/mds-cli/br.cc +++ b/src/tools/mds-cli/br.cc @@ -366,6 +366,8 @@ static blockaccess::BlockAccesserSPtr NewBlockAccesser(const S3Info& s3_info) { .sk = s3_info.sk, .endpoint = s3_info.endpoint, .bucket_name = s3_info.bucket_name}; + blockaccess::FillAwsSdkConfigFromGFlags( + &options.s3_options.aws_sdk_config); auto block_accessor = blockaccess::NewShareBlockAccesser(options); auto status = block_accessor->Init(); diff --git a/test/unit/common/test_config_mapper.cc b/test/unit/common/test_config_mapper.cc index 69020f725..c9ba16e03 100644 --- a/test/unit/common/test_config_mapper.cc +++ b/test/unit/common/test_config_mapper.cc @@ -16,6 +16,10 @@ #include "common/config_mapper.h" +#include + +#include + #include namespace dingofs { @@ -37,6 +41,11 @@ TEST(ConfigMapperTest, FillsS3OptionsFromS3FsInfo) { EXPECT_EQ(options.s3_options.s3_info.sk, "sk-value"); EXPECT_EQ(options.s3_options.s3_info.endpoint, "http://s3.example.com"); EXPECT_EQ(options.s3_options.s3_info.bucket_name, "my-bucket"); + // Regression: the AWS SDK log prefix must come from the --s3_*/--log_dir + // gflags. If it stays empty the SDK writes /.log. + EXPECT_NE(options.s3_options.aws_sdk_config.log_prefix.find( + fmt::format("aws_sdk_{}_", getpid())), + std::string::npos); } TEST(ConfigMapperTest, FillsRadosOptionsFromRadosFsInfo) { diff --git a/test/unit/mds/common/test_type.cc b/test/unit/mds/common/test_type.cc index 87092afe7..05f7f1857 100644 --- a/test/unit/mds/common/test_type.cc +++ b/test/unit/mds/common/test_type.cc @@ -197,6 +197,210 @@ TEST_F(TypeTest, InoTypeAlias) { EXPECT_EQ(sizeof(ino), sizeof(uint64_t)); } +namespace { + +AttrMutationEntry MakeMutation(uint32_t index, uint64_t delta_version) { + AttrMutationEntry mutation; + mutation.set_ino(1); + mutation.set_index(index); + mutation.set_delta_version(delta_version); + return mutation; +} + +} // namespace + +TEST_F(TypeTest, AttrVersionVecConstructFromBaseVersion) { + AttrVersionVec version_vec(100); + + EXPECT_EQ(version_vec.BaseVersion(), 100); + EXPECT_EQ(version_vec.CompleteVersion(), 100); + EXPECT_EQ(version_vec.total_delta_version, 0); + ASSERT_EQ(version_vec.delta_versions.size(), kDirAttrMutationNum); + for (uint32_t i = 0; i < kDirAttrMutationNum; ++i) { + EXPECT_EQ(version_vec.DeltaVersion(i), 0) << "index=" << i; + } + EXPECT_EQ(version_vec.ToString(), "100 0"); +} + +TEST_F(TypeTest, AttrVersionVecConstructFromAttrWithMutation) { + AttrWithMutation attr_with_mutation; + attr_with_mutation.attr.set_ino(1); + attr_with_mutation.attr.set_version(50); + attr_with_mutation.mutations.push_back(MakeMutation(2, 3)); + attr_with_mutation.mutations.push_back(MakeMutation(5, 7)); + + AttrVersionVec version_vec(attr_with_mutation); + + EXPECT_EQ(version_vec.BaseVersion(), 50); + EXPECT_EQ(version_vec.DeltaVersion(2), 3); + EXPECT_EQ(version_vec.DeltaVersion(5), 7); + EXPECT_EQ(version_vec.DeltaVersion(0), 0); + EXPECT_EQ(version_vec.total_delta_version, 10); + EXPECT_EQ(version_vec.CompleteVersion(), 60); + EXPECT_EQ(version_vec.ToString(), "50 10"); +} + +TEST_F(TypeTest, AttrVersionVecDedupsRepeatedMutationSlots) { + // Retried operations may append the same slot twice; the version must not inflate. + AttrWithMutation attr_with_mutation; + attr_with_mutation.attr.set_ino(1); + attr_with_mutation.attr.set_version(50); + attr_with_mutation.mutations.push_back(MakeMutation(2, 3)); + attr_with_mutation.mutations.push_back(MakeMutation(5, 7)); + attr_with_mutation.mutations.push_back(MakeMutation(2, 3)); + attr_with_mutation.mutations.push_back(MakeMutation(5, 7)); + + EXPECT_EQ(attr_with_mutation.TotalDeltaVersion(), 10); + EXPECT_EQ(attr_with_mutation.ToCompleteAttr().version(), 60); + + AttrVersionVec version_vec(attr_with_mutation); + EXPECT_EQ(version_vec.total_delta_version, 10); + EXPECT_EQ(version_vec.CompleteVersion(), 60); +} + +TEST_F(TypeTest, AttrVersionVecPutIfAttrMutationEntry) { + AttrVersionVec version_vec(10); + + // newer delta is accepted and total is adjusted + EXPECT_TRUE(version_vec.PutIf(MakeMutation(1, 5))); + EXPECT_EQ(version_vec.DeltaVersion(1), 5); + EXPECT_EQ(version_vec.total_delta_version, 5); + EXPECT_EQ(version_vec.CompleteVersion(), 15); + + // equal or older delta is rejected + EXPECT_FALSE(version_vec.PutIf(MakeMutation(1, 5))); + EXPECT_FALSE(version_vec.PutIf(MakeMutation(1, 3))); + EXPECT_EQ(version_vec.DeltaVersion(1), 5); + + // larger delta only bumps the total by the difference + EXPECT_TRUE(version_vec.PutIf(MakeMutation(1, 8))); + EXPECT_EQ(version_vec.DeltaVersion(1), 8); + EXPECT_EQ(version_vec.total_delta_version, 8); + + // independent index accumulates + EXPECT_TRUE(version_vec.PutIf(MakeMutation(2, 4))); + EXPECT_EQ(version_vec.total_delta_version, 12); +} + +TEST_F(TypeTest, AttrVersionVecPutIfAttrVersion) { + AttrVersionVec version_vec(10); + + // base version + EXPECT_TRUE(version_vec.PutIf(AttrVersion(uint64_t{20}))); + EXPECT_EQ(version_vec.BaseVersion(), 20); + EXPECT_EQ(version_vec.total_delta_version, 0); + EXPECT_FALSE(version_vec.PutIf(AttrVersion(uint64_t{20}))); + EXPECT_FALSE(version_vec.PutIf(AttrVersion(uint64_t{15}))); + + // delta version + EXPECT_TRUE(version_vec.PutIf(AttrVersion(1, 4))); + EXPECT_EQ(version_vec.DeltaVersion(1), 4); + EXPECT_EQ(version_vec.total_delta_version, 4); + EXPECT_FALSE(version_vec.PutIf(AttrVersion(1, 4))); + EXPECT_FALSE(version_vec.PutIf(AttrVersion(1, 2))); + EXPECT_TRUE(version_vec.PutIf(AttrVersion(1, 6))); + EXPECT_EQ(version_vec.total_delta_version, 6); +} + +TEST_F(TypeTest, AttrVersionVecPutIfAttrEntry) { + AttrVersionVec version_vec(10); + + AttrEntry attr; + attr.set_ino(1); + attr.set_version(20); + + EXPECT_TRUE(version_vec.PutIf(attr)); + EXPECT_EQ(version_vec.BaseVersion(), 20); + EXPECT_FALSE(version_vec.PutIf(attr)); + + attr.set_version(30); + EXPECT_TRUE(version_vec.PutIf(attr)); + EXPECT_EQ(version_vec.BaseVersion(), 30); +} + +TEST_F(TypeTest, AttrVersionVecPutIfAttrWithMutation) { + AttrVersionVec version_vec(10); + + AttrWithMutation attr_with_mutation; + attr_with_mutation.attr.set_ino(1); + attr_with_mutation.attr.set_version(20); + attr_with_mutation.mutations.push_back(MakeMutation(1, 3)); + attr_with_mutation.mutations.push_back(MakeMutation(2, 4)); + + EXPECT_TRUE(version_vec.PutIf(attr_with_mutation)); + EXPECT_EQ(version_vec.BaseVersion(), 20); + EXPECT_EQ(version_vec.DeltaVersion(1), 3); + EXPECT_EQ(version_vec.DeltaVersion(2), 4); + EXPECT_EQ(version_vec.CompleteVersion(), 27); + + // replaying the same attr/mutations is a no-op + EXPECT_FALSE(version_vec.PutIf(attr_with_mutation)); +} + +TEST_F(TypeTest, AttrVersionVecPutIfAttrVersionVec) { + AttrVersionVec version_vec(10); + + AttrVersionVec other(20); + other.PutIf(AttrVersion(1, 5)); + other.PutIf(AttrVersion(2, 7)); + + EXPECT_TRUE(version_vec.PutIf(other)); + EXPECT_EQ(version_vec.BaseVersion(), 20); + EXPECT_EQ(version_vec.DeltaVersion(1), 5); + EXPECT_EQ(version_vec.DeltaVersion(2), 7); + EXPECT_EQ(version_vec.total_delta_version, 12); + + // replaying the same vec is a no-op + EXPECT_FALSE(version_vec.PutIf(other)); + + // base-only newer vec keeps existing deltas + AttrVersionVec base_only(30); + EXPECT_TRUE(version_vec.PutIf(base_only)); + EXPECT_EQ(version_vec.BaseVersion(), 30); + EXPECT_EQ(version_vec.total_delta_version, 12); + + // older base and older deltas are rejected + AttrVersionVec older(5); + EXPECT_FALSE(version_vec.PutIf(older)); + + // newer base but some stale/some newer deltas: only newer deltas win + AttrVersionVec mixed(40); + mixed.PutIf(AttrVersion(1, 2)); // stale -> not applied + mixed.PutIf(AttrVersion(3, 9)); // newer -> applied + EXPECT_TRUE(version_vec.PutIf(mixed)); + EXPECT_EQ(version_vec.BaseVersion(), 40); + EXPECT_EQ(version_vec.DeltaVersion(1), 5); + EXPECT_EQ(version_vec.DeltaVersion(2), 7); + EXPECT_EQ(version_vec.DeltaVersion(3), 9); + EXPECT_EQ(version_vec.total_delta_version, 21); +} + +TEST_F(TypeTest, AttrVersionVecLessThanOrEqual) { + AttrVersionVec version_vec(10); + version_vec.PutIf(AttrVersion(1, 4)); + + EXPECT_TRUE(version_vec.LessThanOrEqual(AttrVersion(uint64_t{10}))); + EXPECT_TRUE(version_vec.LessThanOrEqual(AttrVersion(uint64_t{11}))); + EXPECT_FALSE(version_vec.LessThanOrEqual(AttrVersion(uint64_t{9}))); + + EXPECT_TRUE(version_vec.LessThanOrEqual(AttrVersion(1, 4))); + EXPECT_TRUE(version_vec.LessThanOrEqual(AttrVersion(1, 5))); + EXPECT_FALSE(version_vec.LessThanOrEqual(AttrVersion(1, 3))); +} + +TEST_F(TypeTest, AttrVersionVecGreaterThanOrEqual) { + AttrVersionVec version_vec(10); + version_vec.PutIf(AttrVersion(1, 4)); + + EXPECT_TRUE(version_vec.GreaterThanOrEqual(AttrVersion(uint64_t{10}))); + EXPECT_TRUE(version_vec.GreaterThanOrEqual(AttrVersion(uint64_t{9}))); + EXPECT_FALSE(version_vec.GreaterThanOrEqual(AttrVersion(uint64_t{11}))); + + EXPECT_TRUE(version_vec.GreaterThanOrEqual(AttrVersion(1, 4))); + EXPECT_TRUE(version_vec.GreaterThanOrEqual(AttrVersion(1, 3))); + EXPECT_FALSE(version_vec.GreaterThanOrEqual(AttrVersion(1, 5))); +} + } // namespace unit_test } // namespace mds } // namespace dingofs diff --git a/test/unit/mds/filesystem/test_filesystem_state.cc b/test/unit/mds/filesystem/test_filesystem_state.cc new file mode 100644 index 000000000..bdcc8b81c --- /dev/null +++ b/test/unit/mds/filesystem/test_filesystem_state.cc @@ -0,0 +1,1424 @@ +// Copyright (c) 2023 dingodb.com, Inc. All Rights Reserved +// +// Licensed under the Apache License, Version 2.0 (the "License"); +// you may not use this file except in compliance with the License. +// You may obtain a copy of the License at +// +// http://www.apache.org/licenses/LICENSE-2.0 +// +// Unless required by applicable law or agreed to in writing, software +// distributed under the License is distributed on an "AS IS" BASIS, +// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +// See the License for the specific language governing permissions and +// limitations under the License. + +// Internal-state tests for FileSystem: after each namespace operation, verify +// that the Inode, ShardPartition and DirShard attribute version vectors +// (AttrVersionVec) advance exactly as expected and stay mutually consistent. +// +// The fixture builds a fresh DummyStorage + FileSystem for every test so that +// versions start from 1 and no state leaks between cases (unlike the shared +// static fs in test_filesystem.cc). +// +// Invariants asserted (see review notes): monotonic version, inode==partition +// when the partition is cached, shard never ahead of its partition, fresh +// inode version, read paths do not touch versions, exact serial deltas, and +// the concurrent mutation path. + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "dingofs/error.pb.h" +#include "dingofs/mds.pb.h" +#include "fmt/format.h" +#include "gtest/gtest.h" +#include "json/value.h" +#include "mds/common/context.h" +#include "mds/common/runnable.h" +#include "mds/filesystem/filesystem.h" +#include "mds/filesystem/id_generator.h" +#include "mds/filesystem/store_operation.h" +#include "mds/storage/dummy_storage.h" + +namespace dingofs { +namespace mds { +namespace unit_test { + +namespace { + +constexpr uint32_t kTestFsId = 1000; +constexpr uint64_t kTestMdsId = 1000; +constexpr Ino kTestRootIno = 1; + +pb::mds::FsInfo MakeFsInfo() { + pb::mds::FsInfo fs_info; + fs_info.set_fs_id(kTestFsId); + fs_info.set_fs_name("test_state_fs"); + fs_info.set_fs_type(pb::mds::FsType::S3); + fs_info.set_status(pb::mds::FsStatus::NORMAL); + fs_info.set_block_size(1024 * 1024); + fs_info.set_chunk_size(1024 * 1024 * 64); + fs_info.set_enable_dir_stats(false); + fs_info.set_owner("test"); + fs_info.set_capacity(1024ULL * 1024 * 1024); + fs_info.set_recycle_time_hour(24); + auto* s3_info = fs_info.mutable_extra()->mutable_s3_info(); + s3_info->set_ak("ak"); + s3_info->set_sk("sk"); + s3_info->set_endpoint("http://s3.com"); + s3_info->set_bucketname("bucket"); + + auto* partition_policy = fs_info.mutable_partition_policy(); + partition_policy->set_type(pb::mds::PartitionType::MONOLITHIC_PARTITION); + partition_policy->mutable_mono()->set_mds_id(kTestMdsId); + + return fs_info; +} + +} // namespace + +class FileSystemStateTest : public testing::Test { + protected: + void SetUp() override { + auto storage = DummyStorage::New(); + ASSERT_TRUE(storage->Init("")) << "init kv storage fail."; + storage_ = storage; + + auto processor = OperationProcessor::New(storage_); + ASSERT_TRUE(processor->Init()) << "init operation processor fail."; + processor_ = processor; + + auto inode_id_generator = NewInodeIdGenerator(kTestFsId, storage_); + ASSERT_TRUE(inode_id_generator->Init()) << "init inode id generator fail."; + + auto slice_id_generator = NewSliceIdGenerator(storage_); + ASSERT_TRUE(slice_id_generator->Init()) << "init slice id generator fail."; + + auto worker_set = SimpleWorkerSet::New("fs_state_test", 1, 1024, false, + /*is_inplace_run=*/true); + ASSERT_TRUE(worker_set->Init()) << "init worker set fail."; + worker_set_ = worker_set; + + fs_ = FileSystem::New(kTestMdsId, FsInfo::New(MakeFsInfo()), + std::move(inode_id_generator), slice_id_generator, + processor_, nullptr, nullptr, worker_set_, + worker_set_, nullptr); + ASSERT_TRUE(fs_->CreateRoot().ok()) << "create root fail."; + } + + void TearDown() override { + fs_ = nullptr; + if (processor_ != nullptr) { + processor_->Stop(); + processor_ = nullptr; + } + } + + // --- state accessors --- + InodeSPtr Inode(Ino ino) { return fs_->GetInodeCache().Get(ino); } + PartitionPtr Partition(Ino ino) { return fs_->GetPartitionCache().Get(ino); } + + // Assert that the cached inode and its partition carry identical version + // vectors (base + every delta bucket). + void ExpectInodePartitionVersionEqual(Ino ino) { + auto inode = Inode(ino); + auto partition = Partition(ino); + ASSERT_TRUE(inode != nullptr) << "inode " << ino << " not cached"; + ASSERT_TRUE(partition != nullptr) << "partition " << ino << " not cached"; + + const auto iv = inode->VersionVec(); + const auto& pv = partition->VersionVec(); + EXPECT_EQ(iv.BaseVersion(), pv.BaseVersion()) << "ino=" << ino; + EXPECT_EQ(iv.CompleteVersion(), pv.CompleteVersion()) << "ino=" << ino; + ASSERT_EQ(iv.delta_versions.size(), pv.delta_versions.size()) + << "ino=" << ino; + for (size_t i = 0; i < iv.delta_versions.size(); ++i) { + EXPECT_EQ(iv.DeltaVersion(static_cast(i)), + pv.DeltaVersion(static_cast(i))) + << "ino=" << ino << " delta_index=" << i; + } + } + + // Load the partition for `dir` (and its shards) through a public read path. + void LoadPartition(Ino dir) { + Context ctx; + std::vector entries; + ASSERT_TRUE(fs_->ReadDir(ctx, dir, "", 100, false, entries).ok()) + << "readdir fail for ino " << dir; + ASSERT_TRUE(Partition(dir) != nullptr) + << "partition " << dir << " still not cached"; + } + + // --- operation helpers (each uses its own Context) --- + Ino MkNod(Ino parent, const std::string& name) { + Context ctx; + FileSystem::MkNodParam param; + param.parent = parent; + param.name = name; + param.mode = 0777; + param.uid = 1; + param.gid = 1; + + EntryWithPaOut out; + auto status = fs_->MkNod(ctx, param, out); + EXPECT_TRUE(status.ok()) + << "mknod " << name << " fail: " << status.error_str(); + return out.attr.ino(); + } + + Ino MkDir(Ino parent, const std::string& name) { + Context ctx; + FileSystem::MkDirParam param; + param.parent = parent; + param.name = name; + param.mode = 0777; + param.uid = 1; + param.gid = 1; + + EntryWithPaOut out; + auto status = fs_->MkDir(ctx, param, out); + EXPECT_TRUE(status.ok()) + << "mkdir " << name << " fail: " << status.error_str(); + return out.attr.ino(); + } + + Ino Symlink(Ino parent, const std::string& name, const std::string& target) { + Context ctx; + EntryWithPaOut out; + auto status = fs_->Symlink(ctx, target, parent, name, 1, 1, out); + EXPECT_TRUE(status.ok()) + << "symlink " << name << " fail: " << status.error_str(); + return out.attr.ino(); + } + + // --- best-effort, thread-safe operation helpers (no gtest assertions, + // callers may invoke them from worker threads) --- + // + // The MDS retries store conflicts internally (bounded by + // FLAGS_mds_txn_max_retry_times); a real client retries the op once that + // budget is exhausted. Under this test's heavy same-inode contention the + // budget can run out, so mirror the client and retry ESTORE_MAYBE_RETRY here. + // Any other error (EEXISTED/ENOT_FOUND/...) is returned as-is. + template + bool RetryOnConflict(Fn&& fn) { + for (int attempt = 0; attempt < 200; ++attempt) { + auto status = fn(); + if (status.ok()) return true; + if (status.error_code() != pb::error::ESTORE_MAYBE_RETRY) return false; + std::this_thread::yield(); + } + return false; + } + + bool TryMkNod(Ino parent, const std::string& name) { + return RetryOnConflict([&] { + Context ctx; + FileSystem::MkNodParam param; + param.parent = parent; + param.name = name; + param.mode = 0777; + param.uid = 1; + param.gid = 1; + EntryWithPaOut out; + return fs_->MkNod(ctx, param, out); + }); + } + + bool TryMkDir(Ino parent, const std::string& name) { + return RetryOnConflict([&] { + Context ctx; + FileSystem::MkDirParam param; + param.parent = parent; + param.name = name; + param.mode = 0777; + param.uid = 1; + param.gid = 1; + EntryWithPaOut out; + return fs_->MkDir(ctx, param, out); + }); + } + + bool TryUnLink(Ino parent, const std::string& name) { + return RetryOnConflict([&] { + Context ctx; + EntryWithPaOut out; + return fs_->UnLink(ctx, parent, name, out); + }); + } + + bool TryRmDir(Ino parent, const std::string& name) { + return RetryOnConflict([&] { + Context ctx; + EntryWithPaOut out; + return fs_->RmDir(ctx, parent, name, out); + }); + } + + bool TrySetAttrMode(Ino ino, uint32_t mode) { + return RetryOnConflict([&] { + Context ctx; + FileSystem::SetAttrParam param; + param.to_set = kSetAttrMode; + param.attr.set_fs_id(kTestFsId); + param.attr.set_ino(ino); + param.attr.set_mode(mode); + EntryWithChunkOut out; + return fs_->SetAttr(ctx, ino, param, out); + }); + } + + // --- status-returning helpers (keep the error code for exact checks) --- + Status LookupRaw(Ino parent, const std::string& name, EntryOut& out) { + Context ctx; + return fs_->Lookup(ctx, parent, name, out); + } + + Status BatchCreateOne(Ino parent, const std::string& name, + EntriesWithPaOut& out) { + Context ctx; + FileSystem::MkNodParam param; + param.parent = parent; + param.name = name; + param.mode = 0777; + param.uid = 1; + param.gid = 1; + return fs_->BatchCreate(ctx, parent, {param}, out); + } + + Status MkDirRaw(Ino parent, const std::string& name, EntryWithPaOut& out) { + Context ctx; + FileSystem::MkDirParam param; + param.parent = parent; + param.name = name; + param.mode = 0777; + param.uid = 1; + param.gid = 1; + return fs_->MkDir(ctx, param, out); + } + + Status UnLinkRaw(Ino parent, const std::string& name, EntryWithPaOut& out) { + Context ctx; + return fs_->UnLink(ctx, parent, name, out); + } + + Status RmDirRaw(Ino parent, const std::string& name, EntryWithPaOut& out) { + Context ctx; + return fs_->RmDir(ctx, parent, name, out); + } + + // Read every child name of `dir` through the paged ReadDir path. + std::set SnapshotNames(Ino dir) { + std::set names; + std::string last; + while (true) { + Context ctx; + std::vector entries; + auto status = fs_->ReadDir(ctx, dir, last, 1000, false, entries); + if (!status.ok()) break; + for (const auto& entry : entries) names.insert(entry.name); + if (entries.size() < 1000) break; + last = entries.back().name; + } + return names; + } + + KVStorageSPtr storage_; + OperationProcessorSPtr processor_; + WorkerSetSPtr worker_set_; + FileSystemSPtr fs_; +}; + +// --------------------------------------------------------------------------- +// T1: version/state invariants for namespace write operations +// --------------------------------------------------------------------------- + +TEST_F(FileSystemStateTest, CreateRootState) { + auto inode = Inode(kTestRootIno); + ASSERT_TRUE(inode != nullptr); + EXPECT_EQ(pb::mds::FileType::DIRECTORY, inode->Type()); + EXPECT_EQ(1u, inode->BaseVersion()); + EXPECT_EQ(1u, inode->CompleteVersion()); + + auto partition = Partition(kTestRootIno); + ASSERT_TRUE(partition != nullptr); + EXPECT_EQ("1 0", partition->VersionVec().ToString()); + ExpectInodePartitionVersionEqual(kTestRootIno); +} + +TEST_F(FileSystemStateTest, MkNodAdvancesParentExactly) { + const uint64_t nlink0 = Inode(kTestRootIno)->Nlink(); + const uint64_t base0 = Inode(kTestRootIno)->BaseVersion(); + + for (int i = 0; i < 5; ++i) { + Ino ino = MkNod(kTestRootIno, fmt::format("mn_file_{}", i)); + ASSERT_GT(ino, 0u); + + auto child = Inode(ino); + ASSERT_TRUE(child != nullptr); + EXPECT_EQ(pb::mds::FileType::FILE, child->Type()); + EXPECT_EQ(1u, child->Nlink()); + EXPECT_EQ(1u, child->BaseVersion()); + EXPECT_EQ(1u, child->CompleteVersion()); + + auto root = Inode(kTestRootIno); + EXPECT_EQ(base0 + i + 1, root->BaseVersion()); + EXPECT_EQ(base0 + i + 1, root->CompleteVersion()); + ExpectInodePartitionVersionEqual(kTestRootIno); + + Dentry dentry; + EXPECT_TRUE(Partition(kTestRootIno) + ->Get(fmt::format("mn_file_{}", i), dentry) + .ok()); + } + + // creating a file does not change the parent directory's link count + EXPECT_EQ(nlink0, Inode(kTestRootIno)->Nlink()); +} + +TEST_F(FileSystemStateTest, MkDirBumpsParentNlinkAndVersion) { + const uint64_t nlink0 = Inode(kTestRootIno)->Nlink(); + const uint64_t base0 = Inode(kTestRootIno)->BaseVersion(); + + Ino dir = MkDir(kTestRootIno, "mk_dir_a"); + ASSERT_GT(dir, 0u); + + auto child = Inode(dir); + ASSERT_TRUE(child != nullptr); + EXPECT_EQ(pb::mds::FileType::DIRECTORY, child->Type()); + EXPECT_EQ(static_cast(kEmptyDirMinLinkNum), child->Nlink()); + EXPECT_EQ(1u, child->BaseVersion()); + EXPECT_EQ(1u, child->CompleteVersion()); + + auto root = Inode(kTestRootIno); + EXPECT_EQ(nlink0 + 1, root->Nlink()); + EXPECT_EQ(base0 + 1, root->BaseVersion()); + ExpectInodePartitionVersionEqual(kTestRootIno); +} + +TEST_F(FileSystemStateTest, BatchMkNodAdvancesParentOnce) { + const uint64_t base0 = Inode(kTestRootIno)->BaseVersion(); + + std::vector params; + for (int i = 0; i < 3; ++i) { + FileSystem::MkNodParam param; + param.parent = kTestRootIno; + param.name = fmt::format("bmn_file_{}", i); + param.mode = 0777; + params.push_back(param); + } + + Context ctx; + EntriesWithPaOut out; + ASSERT_TRUE(fs_->BatchMkNod(ctx, params, out).ok()); + ASSERT_EQ(3u, out.attrs.size()); + for (const auto& attr : out.attrs) { + EXPECT_EQ(1u, attr.version()); + auto child = Inode(attr.ino()); + ASSERT_TRUE(child != nullptr); + EXPECT_EQ(1u, child->CompleteVersion()); + } + + // a single batch bumps the parent once, not once per entry + auto root = Inode(kTestRootIno); + EXPECT_EQ(base0 + 1, root->BaseVersion()); + EXPECT_EQ(base0 + 1, root->CompleteVersion()); + ExpectInodePartitionVersionEqual(kTestRootIno); +} + +TEST_F(FileSystemStateTest, BatchMkDirAdvancesParentOnceAndNlinkByN) { + const uint64_t nlink0 = Inode(kTestRootIno)->Nlink(); + const uint64_t base0 = Inode(kTestRootIno)->BaseVersion(); + + std::vector params; + for (int i = 0; i < 4; ++i) { + FileSystem::MkDirParam param; + param.parent = kTestRootIno; + param.name = fmt::format("bmd_dir_{}", i); + param.mode = 0777; + params.push_back(param); + } + + Context ctx; + EntriesWithPaOut out; + ASSERT_TRUE(fs_->BatchMkDir(ctx, params, out).ok()); + ASSERT_EQ(4u, out.attrs.size()); + + auto root = Inode(kTestRootIno); + EXPECT_EQ(nlink0 + 4, root->Nlink()); + EXPECT_EQ(base0 + 1, root->BaseVersion()); + ExpectInodePartitionVersionEqual(kTestRootIno); +} + +TEST_F(FileSystemStateTest, RmDirBumpsParentAndDropsChild) { + Ino dir = MkDir(kTestRootIno, "rm_dir_a"); + ASSERT_TRUE(Partition(kTestRootIno) != nullptr); + + const uint64_t nlink0 = Inode(kTestRootIno)->Nlink(); + const uint64_t base0 = Inode(kTestRootIno)->BaseVersion(); + + Context ctx; + EntryWithPaOut out; + ASSERT_TRUE(fs_->RmDir(ctx, kTestRootIno, "rm_dir_a", out).ok()); + + auto root = Inode(kTestRootIno); + EXPECT_EQ(nlink0 - 1, root->Nlink()); + EXPECT_EQ(base0 + 1, root->BaseVersion()); + ExpectInodePartitionVersionEqual(kTestRootIno); + + EXPECT_TRUE(Inode(dir) == nullptr) + << "removed dir inode should leave the cache"; + Dentry dentry; + EXPECT_FALSE(Partition(kTestRootIno)->Get("rm_dir_a", dentry).ok()); +} + +TEST_F(FileSystemStateTest, LinkAndUnLinkTrackNlinkVersionAndParents) { + Ino dir = MkDir(kTestRootIno, "link_dir_a"); + LoadPartition(dir); + Ino file = MkNod(kTestRootIno, "link_file_a"); + + const uint64_t root_base0 = Inode(kTestRootIno)->BaseVersion(); + const uint64_t dir_base0 = Inode(dir)->BaseVersion(); + const uint64_t file_base0 = Inode(file)->BaseVersion(); + const uint64_t file_nlink0 = Inode(file)->Nlink(); + + { + Context ctx; + EntryWithPaOut out; + ASSERT_TRUE(fs_->Link(ctx, file, dir, "link_child", out).ok()); + + auto child = Inode(file); + ASSERT_TRUE(child != nullptr); + EXPECT_EQ(file_nlink0 + 1, child->Nlink()); + EXPECT_EQ(file_base0 + 1, child->BaseVersion()); + EXPECT_EQ(2u, child->Parents().size()); + + EXPECT_EQ(dir_base0 + 1, Inode(dir)->BaseVersion()); + EXPECT_EQ(root_base0, Inode(kTestRootIno)->BaseVersion()); + ExpectInodePartitionVersionEqual(dir); + } + + { + // remove one of the two links: inode survives with nlink-1 + Context ctx; + EntryWithPaOut out; + ASSERT_TRUE(fs_->UnLink(ctx, dir, "link_child", out).ok()); + + auto child = Inode(file); + ASSERT_TRUE(child != nullptr); + EXPECT_EQ(file_nlink0, child->Nlink()); + EXPECT_FALSE(child->IsDeleted()); + EXPECT_EQ(1u, child->Parents().size()); + EXPECT_EQ(dir_base0 + 2, Inode(dir)->BaseVersion()); + ExpectInodePartitionVersionEqual(dir); + } + + { + // remove the last link: inode becomes deleted + Context ctx; + EntryWithPaOut out; + ASSERT_TRUE(fs_->UnLink(ctx, kTestRootIno, "link_file_a", out).ok()); + + auto child = Inode(file); + ASSERT_TRUE(child != nullptr); + EXPECT_EQ(0u, child->Nlink()); + EXPECT_TRUE(child->IsDeleted()); + } +} + +TEST_F(FileSystemStateTest, SymlinkBumpsParentVersionNotNlink) { + Ino dir = MkDir(kTestRootIno, "sym_dir_a"); + LoadPartition(dir); + + const uint64_t base0 = Inode(dir)->BaseVersion(); + const uint64_t nlink0 = Inode(dir)->Nlink(); + + Ino link = Symlink(dir, "sym_link_a", "/some/target"); + ASSERT_GT(link, 0u); + + auto child = Inode(link); + ASSERT_TRUE(child != nullptr); + EXPECT_EQ(pb::mds::FileType::SYM_LINK, child->Type()); + EXPECT_EQ(1u, child->BaseVersion()); + EXPECT_EQ(1u, child->CompleteVersion()); + EXPECT_EQ("/some/target", child->Symlink()); + + EXPECT_EQ(base0 + 1, Inode(dir)->BaseVersion()); + EXPECT_EQ(nlink0, Inode(dir)->Nlink()); + ExpectInodePartitionVersionEqual(dir); +} + +TEST_F(FileSystemStateTest, SetAttrOnDirRefreshesPartitionVersion) { + Ino dir = MkDir(kTestRootIno, "setattr_dir_a"); + LoadPartition(dir); + + const uint64_t dir_base0 = Inode(dir)->BaseVersion(); + + Context ctx; + FileSystem::SetAttrParam param; + param.to_set = kSetAttrMode; + param.attr.set_fs_id(kTestFsId); + param.attr.set_ino(dir); + param.attr.set_mode(0600); + EntryWithChunkOut out; + ASSERT_TRUE(fs_->SetAttr(ctx, dir, param, out).ok()); + EXPECT_EQ(0600u, out.attr.mode() & 0777); + EXPECT_EQ(dir_base0 + 1, Inode(dir)->BaseVersion()); + ExpectInodePartitionVersionEqual(dir); + + // setattr on a file must not touch its parent directory version + Ino file = MkNod(dir, "setattr_file_a"); + const uint64_t dir_after_mknod = Inode(dir)->BaseVersion(); + + FileSystem::SetAttrParam file_param; + file_param.to_set = kSetAttrMode; + file_param.attr.set_fs_id(kTestFsId); + file_param.attr.set_ino(file); + file_param.attr.set_mode(0644); + EntryWithChunkOut file_out; + ASSERT_TRUE(fs_->SetAttr(ctx, file, file_param, file_out).ok()); + EXPECT_EQ(dir_after_mknod, Inode(dir)->BaseVersion()); +} + +TEST_F(FileSystemStateTest, SetAndRemoveXAttrOnDirRefreshPartitionVersion) { + Ino dir = MkDir(kTestRootIno, "xattr_dir_a"); + LoadPartition(dir); + + const uint64_t base0 = Inode(dir)->BaseVersion(); + + Context ctx; + Inode::XAttrMap xattrs; + xattrs["user.k"] = "v"; + EntryOut out; + ASSERT_TRUE(fs_->SetXAttr(ctx, dir, xattrs, out).ok()); + EXPECT_EQ(base0 + 1, Inode(dir)->BaseVersion()); + ExpectInodePartitionVersionEqual(dir); + + std::string value; + ASSERT_TRUE(fs_->GetXAttr(ctx, dir, "user.k", value).ok()); + EXPECT_EQ("v", value); + + EntryOut remove_out; + ASSERT_TRUE(fs_->RemoveXAttr(ctx, dir, "user.k", remove_out).ok()); + EXPECT_EQ(base0 + 2, Inode(dir)->BaseVersion()); + EXPECT_TRUE(Inode(dir)->XAttr("user.k").empty()); + ExpectInodePartitionVersionEqual(dir); +} + +TEST_F(FileSystemStateTest, RenameSameDirAdvancesParentOnce) { + Ino dir = MkDir(kTestRootIno, "ren1_dir"); + MkNod(dir, "ren1_src"); + LoadPartition(dir); + + const uint64_t base0 = Inode(dir)->BaseVersion(); + + Context ctx; + FileSystem::RenameParam param; + param.old_parent = dir; + param.old_name = "ren1_src"; + param.new_parent = dir; + param.new_name = "ren1_dst"; + FileSystem::RenameResult out; + ASSERT_TRUE(fs_->Rename(ctx, param, out).ok()); + + EXPECT_EQ(base0 + 1, Inode(dir)->BaseVersion()); + ExpectInodePartitionVersionEqual(dir); + + Dentry dentry; + EXPECT_FALSE(Partition(dir)->Get("ren1_src", dentry).ok()); + EXPECT_TRUE(Partition(dir)->Get("ren1_dst", dentry).ok()); + EXPECT_EQ(dir, dentry.ParentIno()); +} + +TEST_F(FileSystemStateTest, RenameDiffDirAdvancesBothParents) { + Ino src_dir = MkDir(kTestRootIno, "ren2_src"); + Ino dst_dir = MkDir(kTestRootIno, "ren2_dst"); + Ino file = MkNod(src_dir, "ren2_file"); + LoadPartition(src_dir); + LoadPartition(dst_dir); + + const uint64_t src_base0 = Inode(src_dir)->BaseVersion(); + const uint64_t dst_base0 = Inode(dst_dir)->BaseVersion(); + + Context ctx; + FileSystem::RenameParam param; + param.old_parent = src_dir; + param.old_name = "ren2_file"; + param.new_parent = dst_dir; + param.new_name = "ren2_file"; + FileSystem::RenameResult out; + ASSERT_TRUE(fs_->Rename(ctx, param, out).ok()); + + EXPECT_EQ(src_base0 + 1, Inode(src_dir)->BaseVersion()); + EXPECT_EQ(dst_base0 + 1, Inode(dst_dir)->BaseVersion()); + ExpectInodePartitionVersionEqual(src_dir); + ExpectInodePartitionVersionEqual(dst_dir); + + auto child = Inode(file); + ASSERT_TRUE(child != nullptr); + auto parents = child->Parents(); + EXPECT_NE(parents.end(), std::find(parents.begin(), parents.end(), dst_dir)); + + Dentry dentry; + EXPECT_FALSE(Partition(src_dir)->Get("ren2_file", dentry).ok()); + EXPECT_TRUE(Partition(dst_dir)->Get("ren2_file", dentry).ok()); +} + +// --------------------------------------------------------------------------- +// T2: read paths and lazy partition loading +// --------------------------------------------------------------------------- + +TEST_F(FileSystemStateTest, ReadPathsDoNotChangeVersions) { + Ino dir = MkDir(kTestRootIno, "read_dir_a"); + Ino file = MkNod(dir, "read_file_a"); + Ino sym = Symlink(dir, "read_sym_a", "/read/target"); + LoadPartition(dir); + + auto inode_version = [&](Ino ino) { + auto inode = Inode(ino); + EXPECT_TRUE(inode != nullptr) << "inode " << ino << " not cached"; + if (inode == nullptr) return std::make_pair(0, 0); + return std::make_pair(inode->BaseVersion(), inode->CompleteVersion()); + }; + auto partition_version = [&](Ino ino) { + auto partition = Partition(ino); + EXPECT_TRUE(partition != nullptr) << "partition " << ino << " not cached"; + if (partition == nullptr) return std::make_pair(0, 0); + return std::make_pair(partition->VersionVec().BaseVersion(), + partition->VersionVec().CompleteVersion()); + }; + + const std::vector inodes = {kTestRootIno, dir, file, sym}; + std::map> inode_before; + for (Ino ino : inodes) inode_before[ino] = inode_version(ino); + + const std::map> partition_before = { + {kTestRootIno, partition_version(kTestRootIno)}, + {dir, partition_version(dir)}, + }; + + Context ctx; + { + EntryOut out; + ASSERT_TRUE(fs_->Lookup(ctx, kTestRootIno, "read_dir_a", out).ok()); + } + { + EntryOut out; + ASSERT_TRUE(fs_->GetAttr(ctx, file, out).ok()); + } + { + Dentry dentry; + ASSERT_TRUE(fs_->GetDentry(ctx, dir, "read_file_a", dentry).ok()); + } + { + std::vector dentries; + ASSERT_TRUE(fs_->ListDentry(ctx, dir, "", 100, false, dentries).ok()); + } + { + std::vector entries; + ASSERT_TRUE(fs_->ReadDir(ctx, dir, "", 100, false, entries).ok()); + } + { + Inode::XAttrMap xattrs; + ASSERT_TRUE(fs_->GetXAttr(ctx, file, xattrs).ok()); + } + { + std::string link; + ASSERT_TRUE(fs_->ReadLink(ctx, sym, link).ok()); + EXPECT_EQ("/read/target", link); + } + { + std::vector outs; + ASSERT_TRUE(fs_->BatchGetInode(ctx, {file, sym}, outs).ok()); + EXPECT_EQ(2u, outs.size()); + } + { + std::vector xattrs; + ASSERT_TRUE(fs_->BatchGetXAttr(ctx, {file}, xattrs).ok()); + EXPECT_EQ(1u, xattrs.size()); + } + + for (Ino ino : inodes) { + EXPECT_EQ(inode_before[ino], inode_version(ino)) + << "inode " << ino << " version changed"; + } + for (const auto& [ino, before] : partition_before) { + EXPECT_EQ(before, partition_version(ino)) + << "partition " << ino << " version changed"; + } +} + +TEST_F(FileSystemStateTest, ShardVersionEqualsPartitionAtFetch) { + Ino dir = MkDir(kTestRootIno, "shard_dir_a"); + ASSERT_TRUE(Partition(dir) == nullptr) + << "fresh dir partition should not be cached"; + + // create children while the partition is not cached: no shard is loaded, the + // dentries only land in the store and as delta ops + MkNod(dir, "s1"); + MkNod(dir, "s2"); + MkNod(dir, "s3"); + ASSERT_TRUE(Partition(dir) == nullptr) + << "mknod must not cache the child partition"; + + // now load the partition + shard through a read path + LoadPartition(dir); + + ExpectInodePartitionVersionEqual(dir); + + Json::Value value; + ASSERT_TRUE(fs_->DescribePartitionShard(dir, value).ok()); + const std::string partition_version = value["version"].asString(); + EXPECT_EQ(Partition(dir)->VersionVec().ToString(), partition_version); + + bool saw_loaded_shard = false; + for (const auto& shard : value["shards"]) { + if (!shard.isMember("id")) continue; // shard not loaded yet + saw_loaded_shard = true; + // fetched from the same store snapshot: shard version == partition version + EXPECT_EQ(partition_version, shard["version"].asString()); + } + EXPECT_TRUE(saw_loaded_shard) << "no shard was loaded"; + + // a further op advances the partition but the already-loaded shard lags + MkNod(dir, "s4"); + Json::Value value_after; + ASSERT_TRUE(fs_->DescribePartitionShard(dir, value_after).ok()); + const std::string partition_version_after = value_after["version"].asString(); + for (const auto& shard : value_after["shards"]) { + if (!shard.isMember("id")) continue; + EXPECT_NE(partition_version_after, shard["version"].asString()); + } +} + +TEST_F(FileSystemStateTest, UncachedPartitionOnlyAdvancesInode) { + Ino dir = MkDir(kTestRootIno, "uncached_dir_a"); + ASSERT_TRUE(Partition(dir) == nullptr); + + const uint64_t base0 = Inode(dir)->BaseVersion(); + MkNod(dir, "u1"); + + // partition is still absent; only the inode side advanced + EXPECT_TRUE(Partition(dir) == nullptr); + EXPECT_EQ(base0 + 1, Inode(dir)->BaseVersion()); + + // loading the partition reconciles both sides from the same store snapshot + { + Context ctx; + EntryOut out; + ASSERT_TRUE(fs_->Lookup(ctx, dir, "u1", out).ok()); + } + ASSERT_TRUE(Partition(dir) != nullptr); + ExpectInodePartitionVersionEqual(dir); +} + +// --------------------------------------------------------------------------- +// T4: concurrent operations still keep inode and partition versions in sync +// --------------------------------------------------------------------------- + +TEST_F(FileSystemStateTest, ConcurrentOpsKeepVersionConsistent) { + constexpr int kThreads = 16; + constexpr int kOpsPerThread = 50; + constexpr int kTotal = kThreads * kOpsPerThread; + + std::atomic ready{0}; + std::atomic go{false}; + std::atomic ok_count{0}; + + std::vector threads; + threads.reserve(kThreads); + for (int t = 0; t < kThreads; ++t) { + threads.emplace_back([&, t] { + ready.fetch_add(1); + while (!go.load(std::memory_order_acquire)) { + std::this_thread::yield(); + } + for (int i = 0; i < kOpsPerThread; ++i) { + Context ctx; + FileSystem::MkNodParam param; + param.parent = kTestRootIno; + param.name = fmt::format("cc_{}_{}", t, i); + param.mode = 0777; + EntryWithPaOut out; + if (fs_->MkNod(ctx, param, out).ok()) ok_count.fetch_add(1); + } + }); + } + + while (ready.load(std::memory_order_acquire) < kThreads) + std::this_thread::yield(); + go.store(true, std::memory_order_release); + for (auto& thread : threads) thread.join(); + + ASSERT_EQ(kTotal, ok_count.load()) << "some concurrent mknod failed"; + + // load shards so Size() reflects the persisted dentries + LoadPartition(kTestRootIno); + EXPECT_EQ(static_cast(kTotal), Partition(kTestRootIno)->Size()); + + // inode and partition must agree bucket by bucket after the dust settles + ExpectInodePartitionVersionEqual(kTestRootIno); + + auto iv = Inode(kTestRootIno)->VersionVec(); + // concurrency must have exercised the mutation (delta) path at least once, + // otherwise this test proves nothing about the in-flight merge logic + EXPECT_GT(iv.total_delta_version, 0u) + << "mutation path was never exercised; run on a multi-core host"; + EXPECT_GT(iv.BaseVersion(), 1u); + EXPECT_EQ(iv.BaseVersion() + iv.total_delta_version, iv.CompleteVersion()); +} + +// Long-running mixed-op concurrency with a checker thread that validates the +// state while operations are still in flight. Rename is deliberately left out +// of the mix: it updates the partition before the inode, so it would break the +// in-flight `partition <= inode` assertion (see ShardPartition/FileSystem +// comment in Rename). All other namespace ops publish the inode first. +// +// Duration defaults to 3000ms; override with LONG_TEST_SECONDS= to run +// longer (e.g. LONG_TEST_SECONDS=120 for a soak run). +TEST_F(FileSystemStateTest, LongRunningConcurrentStateConsistency) { + if (getenv("MANUAL_TEST") == nullptr) { + GTEST_SKIP() << "Skip manual test case."; + } + + const uint64_t run_ms = [] { + if (const char* s = getenv("LONG_TEST_SECONDS")) { + return static_cast(std::stoull(s)) * 1000; + } + return static_cast(3000); + }(); + + constexpr int kWorkers = 8; + constexpr int kKeepersPerWorker = 10; + + // A file that must stay present, untouched, for the whole run. + const std::string stable_name = "long_stable_file"; + Ino stable_ino = MkNod(kTestRootIno, stable_name); + ASSERT_GT(stable_ino, 0u); + LoadPartition(kTestRootIno); // load root shard so reads hit the warm path + + std::atomic stop{false}; + std::atomic ok_ops{0}; + std::atomic checker_checks{0}; + std::atomic checker_failures{0}; + + // Periodic in-flight checker. Runs purely read-only paths and only uses + // EXPECT_* (gtest assertions are safe to call from non-main threads). + std::thread checker([&] { + auto parse_version = [](const std::string& s) { + const size_t space = s.find(' '); + return std::make_pair(std::stoull(s.substr(0, space)), + std::stoull(s.substr(space + 1))); + }; + auto inode_vec = [&] { return Inode(kTestRootIno)->VersionVec(); }; + // Version of the cached partition, or nullopt when it has been evicted. + auto partition_version = + [&]() -> std::optional> { + auto partition = Partition(kTestRootIno); + if (partition == nullptr) return std::nullopt; + Json::Value value; + partition->Dump(value); + return parse_version(value["version"].asString()); + }; + + auto last_inode = inode_vec(); + + while (!stop.load(std::memory_order_acquire)) { + std::this_thread::sleep_for(std::chrono::milliseconds(20)); + checker_checks.fetch_add(1); + + // 1. inode version vector is monotonic, bucket by bucket + auto iv = inode_vec(); + if (iv.BaseVersion() < last_inode.BaseVersion() || + iv.CompleteVersion() < last_inode.CompleteVersion()) { + checker_failures.fetch_add(1); + ADD_FAILURE() << "root inode version went backwards: " + << last_inode.ToString() << " -> " << iv.ToString(); + } + for (size_t i = 0; i < iv.delta_versions.size(); ++i) { + if (iv.DeltaVersion(static_cast(i)) < + last_inode.DeltaVersion(static_cast(i))) { + checker_failures.fetch_add(1); + ADD_FAILURE() << "root inode delta bucket " << i << " went backwards"; + } + } + + // 2. the cached partition must never run ahead of the inode. It is + // rebuilt from the store after eviction, which may yield a lower logical + // version than the evicted cache entry, so only this direction holds + // across evictions. Dump the cache entry directly: once evicted, + // DescribePartitionShard would rebuild a base-only view instead. + auto pv = partition_version(); + if (pv.has_value()) { + uint64_t partition_complete = pv->first + pv->second; + if (partition_complete > iv.CompleteVersion()) { + checker_failures.fetch_add(1); + ADD_FAILURE() << "root partition(" << partition_complete + << ") ahead of inode(" << iv.CompleteVersion() << ")"; + } + } + + // 3. the untouched file stays readable with a stable version + Context ctx; + EntryOut out; + auto status = fs_->Lookup(ctx, kTestRootIno, stable_name, out); + if (!status.ok() || out.attr.ino() != stable_ino || + out.attr.version() != 1) { + checker_failures.fetch_add(1); + ADD_FAILURE() << "stable file changed: status=" << status.error_str() + << " ino=" << out.attr.ino() + << " version=" << out.attr.version(); + } + + last_inode = iv; + } + }); + + // Randomly evict the root partition and/or its dir shards so the workers and + // the checker keep re-exercising the cold, store-backed load path. + std::atomic cleaner_rounds{0}; + std::thread cleaner([&] { + std::mt19937 rng(std::random_device{}()); + std::uniform_int_distribution gap_ms(1, 50); + std::uniform_int_distribution pick(0, 1); + while (!stop.load(std::memory_order_acquire)) { + std::this_thread::sleep_for(std::chrono::milliseconds(gap_ms(rng))); + if (pick(rng) == 0) { + fs_->Test_DeletePartitionFromCache(kTestRootIno); + } else if (auto partition = Partition(kTestRootIno); + partition != nullptr) { + partition->TEST_DeleteDirShard(); + } + cleaner_rounds.fetch_add(1); + } + }); + + std::vector workers; + workers.reserve(kWorkers); + for (int w = 0; w < kWorkers; ++w) { + workers.emplace_back([&, w] { + std::string prev; + bool prev_is_dir = false; + uint64_t i = 0; + while (!stop.load(std::memory_order_acquire)) { + const std::string name = fmt::format("lr_{}_{}", w, i); + const bool is_dir = (i % 2) == 0; + const bool created = is_dir ? TryMkDir(kTestRootIno, name) + : TryMkNod(kTestRootIno, name); + if (created) ok_ops.fetch_add(1); + + // extra concurrent traffic on the same parent inode + if (i % 8 == 0 && + TrySetAttrMode(kTestRootIno, 0700 + static_cast(w % 8))) { + ok_ops.fetch_add(1); + } + + if (!prev.empty()) { + const bool removed = prev_is_dir ? TryRmDir(kTestRootIno, prev) + : TryUnLink(kTestRootIno, prev); + if (removed) ok_ops.fetch_add(1); + } + prev = name; + prev_is_dir = is_dir; + ++i; + } + + // leave a clean, known-final set of entries behind + if (!prev.empty()) { + if (prev_is_dir) { + TryRmDir(kTestRootIno, prev); + } else { + TryUnLink(kTestRootIno, prev); + } + } + for (int j = 0; j < kKeepersPerWorker; ++j) { + const std::string keeper = fmt::format("lrk_{}_{}", w, j); + if (TryMkNod(kTestRootIno, keeper)) ok_ops.fetch_add(1); + } + }); + } + + std::this_thread::sleep_for(std::chrono::milliseconds(run_ms)); + stop.store(true, std::memory_order_release); + checker.join(); + for (auto& worker : workers) worker.join(); + cleaner.join(); + + EXPECT_GT(ok_ops.load(), 0); + EXPECT_GT(checker_checks.load(), 0) << "checker never ran"; + EXPECT_GT(cleaner_rounds.load(), 0) << "cleaner never ran"; + EXPECT_EQ(0, checker_failures.load()) << "in-flight state check failed"; + + // quiescent: both sides must have converged and every expected dentry exists. + // The cleaner may have dropped the partition last, so reload it first. + LoadPartition(kTestRootIno); + ExpectInodePartitionVersionEqual(kTestRootIno); + auto iv = Inode(kTestRootIno)->VersionVec(); + EXPECT_GT(iv.total_delta_version, 0u) << "mutation path was never exercised"; + + std::set expected = {stable_name}; + for (int w = 0; w < kWorkers; ++w) { + for (int j = 0; j < kKeepersPerWorker; ++j) { + expected.insert(fmt::format("lrk_{}_{}", w, j)); + } + } + const auto actual = SnapshotNames(kTestRootIno); + EXPECT_EQ(expected, actual) << "final dentry set mismatch"; +} + +// Long-running soak: each worker owns a private name and drives one full +// Lookup(miss) -> create -> Lookup(hit, version==1) -> delete -> Lookup(miss) +// cycle per iteration, while a cleaner thread keeps evicting the root +// partition and its dir shards. Because the name is private every step has a +// single deterministic outcome, so a failed check is a real bug rather than a +// race on the name; all contention is on the shared root inode/partition. +// +// Duration defaults to 60s; LONG_TEST_SECONDS= overrides it and +// LONG_TEST_SECONDS=0 runs until a check fails. The run stops early as soon as +// any check fails (fail -> ADD_FAILURE + fatal flag, every loop bails out). +// +// WARNING: this is a manual/soak test, it is skipped unless MANUAL_TEST is set. +TEST_F(FileSystemStateTest, LongRunningLookupCreateDeleteVersions) { + if (getenv("MANUAL_TEST") == nullptr) { + GTEST_SKIP() << "Skip manual test case."; + } + + const uint64_t run_ms = [] { + if (const char* s = getenv("LONG_TEST_SECONDS")) { + return static_cast(std::stoull(s)) * 1000; + } + return static_cast(60) * 1000; + }(); + + constexpr int kFileWorkers = 4; + constexpr int kDirWorkers = 4; + + std::atomic stop{false}; + std::atomic fatal{false}; + std::atomic created_ops{0}; + std::atomic deleted_ops{0}; + std::atomic cleaner_rounds{0}; + + // A failed check records the gtest failure and wakes every loop to bail out. + auto Fail = [&](const std::string& msg) { + ADD_FAILURE() << msg; + fatal.store(true, std::memory_order_release); + }; + + // Retry the store conflict budget like a real client. When abort_on_stop is + // set (the worker path) the loop bails as soon as the run is winding down, so + // a 200-attempt loop cannot outlive the duration; the cleanup path passes + // false so it still drains leftovers after stop is set. + auto retry_write = [&](auto&& fn, bool abort_on_stop = true) -> Status { + // Seed with a non-OK status: when the loop aborts before the first attempt + // (stop/fatal already set) callers must not take it for success and read + // the untouched output. + Status last = Status(pb::error::ESTORE_MAYBE_RETRY, "retry aborted"); + for (int attempt = 0; attempt < 200; ++attempt) { + if (fatal.load(std::memory_order_acquire)) return last; + if (abort_on_stop && stop.load(std::memory_order_acquire)) return last; + last = fn(); + if (last.ok() || last.error_code() != pb::error::ESTORE_MAYBE_RETRY) + return last; + std::this_thread::yield(); + } + return last; + }; + + // One worker family: is_dir selects MkDir/RmDir + directory expectations, + // otherwise BatchCreate/UnLink + file expectations. + auto run_worker = [&](int w, bool is_dir) { + const std::string name = + is_dir ? fmt::format("ld_{}", w) : fmt::format("lf_{}", w); + + std::optional last_vec; + uint64_t last_parent_version = 0; + + auto root_complete_version = [&]() -> uint64_t { + auto inode = Inode(kTestRootIno); + return (inode != nullptr) ? inode->CompleteVersion() : 0; + }; + + // Root version vector must never go backwards, and the cached partition + // must never run ahead of the inode. The partition is sampled before the + // inode on purpose: an op publishes the inode first, so any partition + // update observed in the meantime is already reflected in the later inode + // sample. + auto check_root_versions = [&]() -> bool { + auto partition = Partition(kTestRootIno); + const bool has_partition = (partition != nullptr); + const uint64_t partition_complete = + has_partition ? partition->VersionVec().CompleteVersion() : 0; + + auto inode = Inode(kTestRootIno); + if (inode == nullptr) { + Fail(fmt::format("[{}] root inode left the cache", name)); + return false; + } + auto vec = inode->VersionVec(); + + if (last_vec.has_value()) { + if (vec.BaseVersion() < last_vec->BaseVersion() || + vec.CompleteVersion() < last_vec->CompleteVersion()) { + Fail(fmt::format("[{}] root version went backwards: {} -> {}", name, + last_vec->ToString(), vec.ToString())); + return false; + } + for (size_t i = 0; i < vec.delta_versions.size(); ++i) { + const auto idx = static_cast(i); + if (vec.DeltaVersion(idx) < last_vec->DeltaVersion(idx)) { + Fail(fmt::format("[{}] root delta bucket {} went backwards", name, + idx)); + return false; + } + } + } + last_vec = vec; + + if (has_partition && partition_complete > vec.CompleteVersion()) { + Fail(fmt::format("[{}] root partition({}) ahead of inode({})", name, + partition_complete, vec.CompleteVersion())); + return false; + } + + return true; + }; + + while (!stop.load(std::memory_order_acquire) && + !fatal.load(std::memory_order_acquire)) { + // 1. the private name must be absent after the previous cycle + { + EntryOut out; + auto status = LookupRaw(kTestRootIno, name, out); + if (status.ok()) { + Fail(fmt::format( + "[{}] step1: expected absent, found ino={} version={}", name, + out.attr.ino(), out.attr.version())); + return; + } + if (status.error_code() != pb::error::ENOT_FOUND) { + Fail(fmt::format("[{}] step1: unexpected lookup error: {}", name, + status.error_str())); + return; + } + } + + const uint64_t parent_before = root_complete_version(); + + // 2. create (batch size one), then verify the fresh inode + Ino ino = 0; + uint64_t parent_after_create = 0; + if (is_dir) { + EntryWithPaOut out; + auto status = + retry_write([&] { return MkDirRaw(kTestRootIno, name, out); }); + if (!status.ok()) { + if (stop.load(std::memory_order_acquire) || + fatal.load(std::memory_order_acquire)) + return; + Fail(fmt::format("[{}] step2: mkdir failed: {}", name, + status.error_str())); + return; + } + ino = out.attr.ino(); + parent_after_create = out.parent_attr.version(); + if (out.attr.version() != 1 || + out.attr.type() != pb::mds::FileType::DIRECTORY || + out.attr.nlink() != static_cast(kEmptyDirMinLinkNum)) { + Fail(fmt::format( + "[{}] step2: bad new dir attr ino={} version={} type={} nlink={}", + name, out.attr.ino(), out.attr.version(), + static_cast(out.attr.type()), out.attr.nlink())); + return; + } + } else { + EntriesWithPaOut out; + auto status = retry_write( + [&] { return BatchCreateOne(kTestRootIno, name, out); }); + if (!status.ok()) { + if (stop.load(std::memory_order_acquire) || + fatal.load(std::memory_order_acquire)) + return; + Fail(fmt::format("[{}] step2: batch create failed: {}", name, + status.error_str())); + return; + } + if (out.attrs.size() != 1) { + Fail(fmt::format("[{}] step2: batch create returned {} attrs", name, + out.attrs.size())); + return; + } + ino = out.attrs[0].ino(); + parent_after_create = out.parent_attr.version(); + if (out.attrs[0].version() != 1 || + out.attrs[0].type() != pb::mds::FileType::FILE || + out.attrs[0].nlink() != 1) { + Fail(fmt::format( + "[{}] step2: bad new file attr ino={} version={} type={} " + "nlink={}", + name, out.attrs[0].ino(), out.attrs[0].version(), + static_cast(out.attrs[0].type()), out.attrs[0].nlink())); + return; + } + } + if (parent_after_create <= parent_before || + parent_after_create <= last_parent_version) { + Fail(fmt::format( + "[{}] step2: parent version did not advance: before={} returned={} " + "prev={}", + name, parent_before, parent_after_create, last_parent_version)); + return; + } + last_parent_version = parent_after_create; + if (!check_root_versions()) return; + created_ops.fetch_add(1, std::memory_order_relaxed); + + // 3. the created inode must be visible at exactly version one + { + EntryOut out; + auto status = LookupRaw(kTestRootIno, name, out); + if (!status.ok() || out.attr.ino() != ino || out.attr.version() != 1) { + Fail( + fmt::format("[{}] step3: lookup after create: status={} ino={} " + "expected_ino={} version={}", + name, status.error_str(), out.attr.ino(), ino, + out.attr.version())); + return; + } + } + + // 4. delete it + { + EntryWithPaOut out; + auto status = retry_write([&] { + return is_dir ? RmDirRaw(kTestRootIno, name, out) + : UnLinkRaw(kTestRootIno, name, out); + }); + if (!status.ok()) { + if (stop.load(std::memory_order_acquire) || + fatal.load(std::memory_order_acquire)) + return; + Fail(fmt::format("[{}] step4: {} failed: {}", name, + is_dir ? "rmdir" : "unlink", status.error_str())); + return; + } + if (out.attr.ino() != ino) { + Fail(fmt::format("[{}] step4: removed ino={} expected={}", name, + out.attr.ino(), ino)); + return; + } + if (out.parent_attr.version() <= last_parent_version) { + Fail(fmt::format( + "[{}] step4: parent version did not advance: returned={} prev={}", + name, out.parent_attr.version(), last_parent_version)); + return; + } + last_parent_version = out.parent_attr.version(); + } + if (!check_root_versions()) return; + deleted_ops.fetch_add(1, std::memory_order_relaxed); + + // 5. it must be gone again, and reads must say exactly so + { + EntryOut out; + auto status = LookupRaw(kTestRootIno, name, out); + if (status.ok()) { + Fail(fmt::format( + "[{}] step5: expected absent after delete, found ino={}", name, + out.attr.ino())); + return; + } + if (status.error_code() != pb::error::ENOT_FOUND) { + Fail(fmt::format("[{}] step5: unexpected lookup error: {}", name, + status.error_str())); + return; + } + } + } + }; + + // Randomly evict the root partition and/or its dir shards so the workers + // keep re-exercising the cold, store-backed load path. + std::thread cleaner([&] { + std::mt19937 rng(std::random_device{}()); + std::uniform_int_distribution gap_ms(1, 50); + std::uniform_int_distribution pick(0, 1); + while (!stop.load(std::memory_order_acquire) && + !fatal.load(std::memory_order_acquire)) { + std::this_thread::sleep_for(std::chrono::milliseconds(gap_ms(rng))); + if (pick(rng) == 0) { + fs_->Test_DeletePartitionFromCache(kTestRootIno); + } else if (auto partition = Partition(kTestRootIno); + partition != nullptr) { + partition->TEST_DeleteDirShard(); + } + cleaner_rounds.fetch_add(1, std::memory_order_relaxed); + } + }); + + std::vector workers; + workers.reserve(kFileWorkers + kDirWorkers); + for (int w = 0; w < kFileWorkers; ++w) + workers.emplace_back([&, w] { run_worker(w, false); }); + for (int w = 0; w < kDirWorkers; ++w) + workers.emplace_back([&, w] { run_worker(w, true); }); + + // LONG_TEST_SECONDS=0 means: run until a check fails. + if (run_ms == 0) { + while (!fatal.load(std::memory_order_acquire)) { + std::this_thread::sleep_for(std::chrono::milliseconds(100)); + } + } else { + const auto deadline = + std::chrono::steady_clock::now() + std::chrono::milliseconds(run_ms); + while (!fatal.load(std::memory_order_acquire) && + std::chrono::steady_clock::now() < deadline) { + std::this_thread::sleep_for(std::chrono::milliseconds(20)); + } + } + stop.store(true, std::memory_order_release); + for (auto& worker : workers) worker.join(); + cleaner.join(); + + // best-effort: a stop can land mid-cycle, so leave a clean root behind + for (int w = 0; w < kFileWorkers; ++w) { + EntryWithPaOut out; + retry_write( + [&] { return UnLinkRaw(kTestRootIno, fmt::format("lf_{}", w), out); }, + false); + } + for (int w = 0; w < kDirWorkers; ++w) { + EntryWithPaOut out; + retry_write( + [&] { return RmDirRaw(kTestRootIno, fmt::format("ld_{}", w), out); }, + false); + } + + EXPECT_FALSE(fatal.load()) << "strict check failed during the run"; + + // quiescent: after the last eviction both sides must have converged + if (!fatal.load()) { + LoadPartition(kTestRootIno); + ExpectInodePartitionVersionEqual(kTestRootIno); + EXPECT_EQ(0u, SnapshotNames(kTestRootIno).size()) + << "leftover entries in root"; + } + + EXPECT_GT(created_ops.load(), 0u) << "worker never created anything"; + EXPECT_GT(deleted_ops.load(), 0u) << "worker never deleted anything"; + EXPECT_GT(cleaner_rounds.load(), 0u) << "cleaner never ran"; +} + +} // namespace unit_test +} // namespace mds +} // namespace dingofs diff --git a/test/unit/mds/filesystem/test_inode.cc b/test/unit/mds/filesystem/test_inode.cc index b7de95ce8..cad01ad13 100644 --- a/test/unit/mds/filesystem/test_inode.cc +++ b/test/unit/mds/filesystem/test_inode.cc @@ -17,6 +17,7 @@ #include #include #include +#include #include #include #include @@ -73,12 +74,47 @@ static pb::mds::Inode GenInode(uint32_t fs_id, uint64_t ino, return inode; } +static AttrMutationEntry GenMutation(uint64_t ino, uint32_t index, + uint64_t delta_version) { + AttrMutationEntry mutation; + mutation.set_ino(ino); + mutation.set_index(index); + mutation.set_delta_version(delta_version); + + auto now_ns = utils::TimestampNs(); + mutation.set_ctime(now_ns); + mutation.set_mtime(now_ns); + mutation.set_atime(now_ns); + + return mutation; +} + +static AttrWithMutation GenAttrWithMutation( + uint32_t fs_id, uint64_t ino, uint64_t base_version, + std::initializer_list> mutations) { + AttrWithMutation attr_with_mutation; + attr_with_mutation.attr = + GenInode(fs_id, ino, pb::mds::FileType::DIRECTORY, base_version); + for (const auto& [index, delta_version] : mutations) { + attr_with_mutation.mutations.push_back( + GenMutation(ino, index, delta_version)); + } + + return attr_with_mutation; +} + class InodeCacheTest : public testing::Test { protected: void SetUp() override {} void TearDown() override {} }; +class InodeTest : public testing::Test { + protected: + void SetUp() override {} + void TearDown() override {} +}; + TEST_F(InodeCacheTest, Put) { InodeCache inode_cache(kFsId); @@ -118,7 +154,7 @@ TEST_F(InodeCacheTest, Put) { ASSERT_EQ(inode->Ino(), ino); ASSERT_EQ(inode->Gid(), 1008); ASSERT_EQ(inode->Uid(), 1008); - ASSERT_EQ(inode->Version(), 1); + ASSERT_EQ(inode->CompleteVersion(), 1); auto attr_entry = GenInode(kFsId, ino, pb::mds::FileType::DIRECTORY); attr_entry.set_gid(1234); @@ -133,7 +169,7 @@ TEST_F(InodeCacheTest, Put) { ASSERT_EQ(inode->Gid(), 1234); ASSERT_EQ(inode->Uid(), 5678); ASSERT_EQ(inode->Length(), 1234567); - ASSERT_EQ(inode->Version(), 2); + ASSERT_EQ(inode->CompleteVersion(), 2); } { @@ -188,7 +224,7 @@ TEST_F(InodeCacheTest, Put) { ASSERT_EQ(inode->Gid(), 1234); ASSERT_EQ(inode->Uid(), 5678); ASSERT_EQ(inode->Length(), 1234567); - ASSERT_EQ(inode->Version(), 2); + ASSERT_EQ(inode->CompleteVersion(), 2); } // put by AttrEntry&& @@ -202,7 +238,7 @@ TEST_F(InodeCacheTest, Put) { auto inode = inode_cache.Get(ino); ASSERT_TRUE(inode != nullptr); ASSERT_EQ(inode->Ino(), ino); - ASSERT_EQ(inode->Version(), 1); + ASSERT_EQ(inode->CompleteVersion(), 1); // update auto attr_entry = GenInode(kFsId, ino, pb::mds::FileType::DIRECTORY); @@ -223,7 +259,7 @@ TEST_F(InodeCacheTest, Put) { ASSERT_EQ(inode->Gid(), 1234); ASSERT_EQ(inode->Uid(), 5678); ASSERT_EQ(inode->Length(), 1234567); - ASSERT_EQ(inode->Version(), 2); + ASSERT_EQ(inode->CompleteVersion(), 2); } } @@ -325,6 +361,108 @@ TEST_F(InodeCacheTest, Get) { ASSERT_EQ(inodes[4]->Ino(), 4005); } +TEST_F(InodeTest, BaseVersionPutIfIsMonotonic) { + const Ino ino = 6001; + auto inode = Inode::New(GenInode(kFsId, ino, pb::mds::FileType::DIRECTORY, 10)); + + ASSERT_EQ(inode->BaseVersion(), 10); + ASSERT_EQ(inode->CompleteVersion(), 10); + + // stale attr is rejected, neither version nor fields change + auto stale = GenInode(kFsId, ino, pb::mds::FileType::DIRECTORY, 5); + stale.set_length(999); + inode->PutIf(stale, "test"); + ASSERT_EQ(inode->BaseVersion(), 10); + ASSERT_EQ(inode->CompleteVersion(), 10); + ASSERT_EQ(inode->Length(), 0); + + // equal version is rejected too + inode->PutIf(GenInode(kFsId, ino, pb::mds::FileType::DIRECTORY, 10), "test"); + ASSERT_EQ(inode->BaseVersion(), 10); + ASSERT_EQ(inode->CompleteVersion(), 10); + + // newer attr wins and updates fields + auto newer = GenInode(kFsId, ino, pb::mds::FileType::DIRECTORY, 12); + newer.set_length(123); + inode->PutIf(newer, "test"); + ASSERT_EQ(inode->BaseVersion(), 12); + ASSERT_EQ(inode->CompleteVersion(), 12); + ASSERT_EQ(inode->Length(), 123); + ASSERT_EQ(inode->ToAttr().version(), 12); +} + +TEST_F(InodeTest, CompleteVersionAccumulatesMutations) { + const Ino ino = 6002; + auto inode = Inode::New(GenInode(kFsId, ino, pb::mds::FileType::DIRECTORY, 100)); + + auto attr_with_mutation = GenAttrWithMutation(kFsId, ino, 110, {{1, 3}, {2, 4}}); + inode->PutIf(attr_with_mutation, "test"); + ASSERT_EQ(inode->BaseVersion(), 110); + ASSERT_EQ(inode->CompleteVersion(), 117); + ASSERT_EQ(inode->ToAttr().version(), 117); + + // replaying the same attr/mutations is a no-op + inode->PutIf(attr_with_mutation, "test"); + ASSERT_EQ(inode->BaseVersion(), 110); + ASSERT_EQ(inode->CompleteVersion(), 117); + + // stale base version but newer delta still advances the delta + inode->PutIf(GenAttrWithMutation(kFsId, ino, 105, {{1, 5}}), "test"); + ASSERT_EQ(inode->BaseVersion(), 110); + ASSERT_EQ(inode->CompleteVersion(), 119); + + // newer base version keeps already accumulated deltas + inode->PutIf(GenInode(kFsId, ino, pb::mds::FileType::DIRECTORY, 120), + "test"); + ASSERT_EQ(inode->BaseVersion(), 120); + ASSERT_EQ(inode->CompleteVersion(), 129); +} + +TEST_F(InodeTest, PutByMutationIsMonotonic) { + const Ino ino = 6003; + auto inode = Inode::New(GenInode(kFsId, ino, pb::mds::FileType::DIRECTORY, 10)); + + auto attr = inode->PutByMutation(GenMutation(ino, 1, 5), "test"); + ASSERT_EQ(attr.version(), 15); + ASSERT_EQ(inode->CompleteVersion(), 15); + + // equal and older delta are rejected, current attr is returned unchanged + attr = inode->PutByMutation(GenMutation(ino, 1, 5), "test"); + ASSERT_EQ(attr.version(), 15); + attr = inode->PutByMutation(GenMutation(ino, 1, 3), "test"); + ASSERT_EQ(attr.version(), 15); + ASSERT_EQ(inode->CompleteVersion(), 15); + + // newer delta on the same index only adds the difference + attr = inode->PutByMutation(GenMutation(ino, 1, 8), "test"); + ASSERT_EQ(attr.version(), 18); + ASSERT_EQ(inode->CompleteVersion(), 18); + + // independent index accumulates + attr = inode->PutByMutation(GenMutation(ino, 2, 4), "test"); + ASSERT_EQ(attr.version(), 22); + ASSERT_EQ(inode->CompleteVersion(), 22); + + // base version bump keeps the mutation deltas + auto newer = GenInode(kFsId, ino, pb::mds::FileType::DIRECTORY, 30); + inode->PutIf(newer, "test"); + ASSERT_EQ(inode->BaseVersion(), 30); + ASSERT_EQ(inode->CompleteVersion(), 42); +} + +TEST_F(InodeTest, ConstructFromAttrWithMutation) { + const Ino ino = 6005; + auto attr_with_mutation = GenAttrWithMutation(kFsId, ino, 50, {{2, 3}, {5, 7}}); + + auto inode = Inode::New(attr_with_mutation); + + ASSERT_EQ(inode->Ino(), ino); + ASSERT_EQ(inode->Type(), pb::mds::FileType::DIRECTORY); + ASSERT_EQ(inode->BaseVersion(), 50); + ASSERT_EQ(inode->CompleteVersion(), 60); + ASSERT_EQ(inode->ToAttr().version(), 60); +} + TEST_F(InodeCacheTest, Benchmark) { if (getenv("SLOW_TEST") == nullptr) { GTEST_SKIP() << "Skip slow test case."; diff --git a/test/unit/mds/filesystem/test_partition.cc b/test/unit/mds/filesystem/test_partition.cc index fb0adb8ae..99254758f 100644 --- a/test/unit/mds/filesystem/test_partition.cc +++ b/test/unit/mds/filesystem/test_partition.cc @@ -15,7 +15,9 @@ #include #include #include +#include #include +#include #include #include "dingofs/mds.pb.h" @@ -85,6 +87,35 @@ static pb::mds::Dentry GenDentry(uint32_t fs_id, uint64_t parent, uint64_t ino, return dentry; } +static AttrMutationEntry GenMutation(uint64_t ino, uint32_t index, + uint64_t delta_version) { + AttrMutationEntry mutation; + mutation.set_ino(ino); + mutation.set_index(index); + mutation.set_delta_version(delta_version); + + auto now_ns = utils::TimestampNs(); + mutation.set_ctime(now_ns); + mutation.set_mtime(now_ns); + mutation.set_atime(now_ns); + + return mutation; +} + +static AttrWithMutation GenAttrWithMutation( + uint32_t fs_id, uint64_t ino, uint64_t base_version, + std::initializer_list> mutations) { + AttrWithMutation attr_with_mutation; + attr_with_mutation.attr = + GenInode(fs_id, ino, pb::mds::FileType::DIRECTORY, base_version); + for (const auto& [index, delta_version] : mutations) { + attr_with_mutation.mutations.push_back( + GenMutation(ino, index, delta_version)); + } + + return attr_with_mutation; +} + // Mock OperationProcessor for testing class MockOperationProcessor : public OperationProcessor { public: @@ -123,7 +154,7 @@ TEST_F(DirShardTest, BasicPutGetDelete) { ASSERT_TRUE(shard != nullptr); ASSERT_EQ(shard->ID(), 1); - ASSERT_EQ(shard->Version(), 1); + ASSERT_EQ(shard->VersionVec().BaseVersion(), 1); ASSERT_TRUE(shard->Empty()); // Put dentry @@ -334,7 +365,61 @@ TEST_F(DirShardTest, ToString) { std::string str = shard->ToString(); ASSERT_NE(str.find("id(1)"), std::string::npos); - ASSERT_NE(str.find("version(1)"), std::string::npos); + ASSERT_NE(str.find("version(1 0)"), std::string::npos); +} + +TEST_F(DirShardTest, VersionVecFromBaseVersion) { + DirShardSPtr shard = + DirShard::New(1, Range{"", ""}, 10, std::vector{}); + + ASSERT_EQ(shard->VersionVec().BaseVersion(), 10); + ASSERT_EQ(shard->VersionVec().CompleteVersion(), 10); + ASSERT_EQ(shard->VersionVec().total_delta_version, 0); + ASSERT_EQ(shard->VersionVec().DeltaVersion(0), 0); + ASSERT_EQ(shard->VersionString(), "10 0"); +} + +TEST_F(DirShardTest, VersionVecFromAttrWithMutation) { + AttrVersionVec version_vec( + GenAttrWithMutation(kFsId, kParentIno, 50, {{2, 3}, {5, 7}})); + + DirShardSPtr shard = + DirShard::New(1, Range{"", ""}, version_vec, std::vector{}); + + ASSERT_EQ(shard->VersionVec().BaseVersion(), 50); + ASSERT_EQ(shard->VersionVec().DeltaVersion(2), 3); + ASSERT_EQ(shard->VersionVec().DeltaVersion(5), 7); + ASSERT_EQ(shard->VersionVec().CompleteVersion(), 60); + ASSERT_EQ(shard->VersionString(), "50 10"); +} + +TEST_F(DirShardTest, SplitPreservesVersionVec) { + AttrVersionVec version_vec(100); + version_vec.PutIf(AttrVersion(1, 5)); + + std::vector dentries; + for (int i = 0; i < 10; ++i) { + dentries.emplace_back(GenDentry(kFsId, kParentIno, 200 + i, + fmt::format("file{:02d}", i), + pb::mds::FileType::FILE)); + } + DirShardSPtr shard = + DirShard::New(1, Range{"", ""}, version_vec, dentries); + + auto [left, right] = shard->Split("file05", 2, 3); + + // both halves inherit the parent shard version + for (const auto& half : {left, right}) { + ASSERT_EQ(half->VersionVec().BaseVersion(), 100); + ASSERT_EQ(half->VersionVec().DeltaVersion(1), 5); + ASSERT_EQ(half->VersionVec().CompleteVersion(), 105); + } + + // version vectors are value copies, mutating the source has no effect + version_vec.PutIf(AttrVersion(1, 9)); + ASSERT_EQ(shard->VersionVec().CompleteVersion(), 105); + ASSERT_EQ(left->VersionVec().CompleteVersion(), 105); + ASSERT_EQ(right->VersionVec().CompleteVersion(), 105); } TEST_F(DirShardTest, UpdateLastActiveTime) { @@ -655,7 +740,7 @@ TEST_F(ShardPartitionBasicTest, BasicProperties) { ASSERT_EQ(partition_->FsId(), kFsId); ASSERT_EQ(partition_->INo(), kParentIno); ASSERT_EQ(partition_->BaseVersion(), 1); - ASSERT_EQ(partition_->DeltaVersion(), 1); + ASSERT_EQ(partition_->CompleteVersion(), 1); } TEST_F(ShardPartitionBasicTest, PutWithVersion) { @@ -756,7 +841,7 @@ TEST_F(ShardPartitionBasicTest, Refresh) { // Create new inode with higher version auto new_inode = Inode::New(GenInode(kFsId, kParentIno, pb::mds::FileType::DIRECTORY, 2)); - ASSERT_EQ(new_inode->Version(), 2); + ASSERT_EQ(new_inode->CompleteVersion(), 2); // Create new partition with higher version inode auto partition2 = ShardPartition::New( @@ -792,24 +877,24 @@ TEST_F(ShardPartitionBasicTest, PartitionCacheIntegration) { } TEST_F(ShardPartitionBasicTest, DeltaVersionTracking) { - ASSERT_EQ(partition_->DeltaVersion(), 1); + ASSERT_EQ(partition_->CompleteVersion(), 1); // Put dentry with version Dentry dentry( GenDentry(kFsId, kParentIno, 200, "file1", pb::mds::FileType::FILE)); partition_->Put(dentry, 3); - ASSERT_EQ(partition_->DeltaVersion(), 3); + ASSERT_EQ(partition_->CompleteVersion(), 3); // Delete with higher version partition_->Delete("file1", 5); - ASSERT_EQ(partition_->DeltaVersion(), 5); + ASSERT_EQ(partition_->CompleteVersion(), 5); // Delete with lower version should not change delta_version partition_->Delete("file1", 4); - ASSERT_EQ(partition_->DeltaVersion(), 5); + ASSERT_EQ(partition_->CompleteVersion(), 5); } TEST_F(ShardPartitionBasicTest, DeleteNonExistent) { @@ -820,6 +905,189 @@ TEST_F(ShardPartitionBasicTest, DeleteNonExistent) { ASSERT_EQ(partition_->Size(), 0); } +class ShardPartitionVersionTest : public testing::Test { + protected: + void SetUp() override { + mock_processor_ = std::make_shared(); + } + + void TearDown() override {} + + PartitionPtr NewPartition(uint64_t version) { + return ShardPartition::New( + mock_processor_, + GenInode(kFsId, kParentIno, pb::mds::FileType::DIRECTORY, version)); + } + + std::shared_ptr mock_processor_; +}; + +TEST_F(ShardPartitionVersionTest, BaseVersionFromAttrEntry) { + auto partition = NewPartition(10); + + ASSERT_EQ(partition->BaseVersion(), 10); + ASSERT_EQ(partition->CompleteVersion(), 10); + ASSERT_EQ(partition->VersionVec().total_delta_version, 0); + ASSERT_EQ(partition->VersionVec().ToString(), "10 0"); +} + +TEST_F(ShardPartitionVersionTest, BaseVersionFromAttrWithMutation) { + auto partition = ShardPartition::New( + mock_processor_, + GenAttrWithMutation(kFsId, kParentIno, 50, {{2, 3}, {5, 7}})); + + ASSERT_EQ(partition->BaseVersion(), 50); + ASSERT_EQ(partition->VersionVec().DeltaVersion(2), 3); + ASSERT_EQ(partition->VersionVec().DeltaVersion(5), 7); + ASSERT_EQ(partition->CompleteVersion(), 60); + ASSERT_EQ(partition->VersionVec().ToString(), "50 10"); +} + +TEST_F(ShardPartitionVersionTest, PutWithBaseVersionIsMonotonic) { + auto partition = NewPartition(10); + Dentry dentry( + GenDentry(kFsId, kParentIno, 200, "file1", pb::mds::FileType::FILE)); + + partition->Put(dentry, 20); + ASSERT_EQ(partition->BaseVersion(), 20); + ASSERT_EQ(partition->CompleteVersion(), 20); + + // stale and equal base versions are rejected + partition->Put(dentry, 15); + ASSERT_EQ(partition->BaseVersion(), 20); + ASSERT_EQ(partition->CompleteVersion(), 20); + + partition->Put(dentry, 20); + ASSERT_EQ(partition->BaseVersion(), 20); + ASSERT_EQ(partition->CompleteVersion(), 20); +} + +TEST_F(ShardPartitionVersionTest, PutWithDeltaVersionAccumulates) { + auto partition = NewPartition(100); + Dentry dentry( + GenDentry(kFsId, kParentIno, 200, "file1", pb::mds::FileType::FILE)); + + partition->Put(dentry, AttrVersion(1, 5)); + ASSERT_EQ(partition->BaseVersion(), 100); + ASSERT_EQ(partition->VersionVec().DeltaVersion(1), 5); + ASSERT_EQ(partition->CompleteVersion(), 105); + + // replaying the same delta is a no-op + partition->Put(dentry, AttrVersion(1, 5)); + ASSERT_EQ(partition->CompleteVersion(), 105); + + // stale delta is rejected + partition->Put(dentry, AttrVersion(1, 3)); + ASSERT_EQ(partition->CompleteVersion(), 105); + + // newer delta on the same index only adds the difference + partition->Put(dentry, AttrVersion(1, 8)); + ASSERT_EQ(partition->CompleteVersion(), 108); + + // independent index accumulates + partition->Put(dentry, AttrVersion(2, 4)); + ASSERT_EQ(partition->VersionVec().DeltaVersion(1), 8); + ASSERT_EQ(partition->VersionVec().DeltaVersion(2), 4); + ASSERT_EQ(partition->CompleteVersion(), 112); +} + +TEST_F(ShardPartitionVersionTest, DeleteWithDeltaVersionIsMonotonic) { + auto partition = NewPartition(100); + + partition->Delete("file1", AttrVersion(1, 5)); + ASSERT_EQ(partition->BaseVersion(), 100); + ASSERT_EQ(partition->CompleteVersion(), 105); + + // stale delta is rejected + partition->Delete("file1", AttrVersion(1, 3)); + ASSERT_EQ(partition->CompleteVersion(), 105); + + // base version bump keeps the accumulated delta + partition->Delete("file1", 200); + ASSERT_EQ(partition->BaseVersion(), 200); + ASSERT_EQ(partition->CompleteVersion(), 205); +} + +TEST_F(ShardPartitionVersionTest, DeleteMultipleWithDeltaVersion) { + auto partition = NewPartition(100); + std::vector names = {"file0", "file1"}; + + partition->Delete(names, AttrVersion(2, 4)); + ASSERT_EQ(partition->BaseVersion(), 100); + ASSERT_EQ(partition->VersionVec().DeltaVersion(2), 4); + ASSERT_EQ(partition->CompleteVersion(), 104); + + // stale delta is rejected + partition->Delete(names, AttrVersion(2, 3)); + ASSERT_EQ(partition->CompleteVersion(), 104); +} + +TEST_F(ShardPartitionVersionTest, RefreshVersionMerges) { + auto partition = NewPartition(10); + + partition->RefreshVersion(AttrVersion(1, 5)); + ASSERT_EQ(partition->BaseVersion(), 10); + ASSERT_EQ(partition->CompleteVersion(), 15); + + // base version bump keeps the delta + partition->RefreshVersion(AttrVersion(uint64_t{20})); + ASSERT_EQ(partition->BaseVersion(), 20); + ASSERT_EQ(partition->CompleteVersion(), 25); + + // stale version is rejected + partition->RefreshVersion(AttrVersion(uint64_t{15})); + ASSERT_EQ(partition->BaseVersion(), 20); + ASSERT_EQ(partition->CompleteVersion(), 25); +} + +TEST_F(ShardPartitionVersionTest, CachePutIfMergesVersionVec) { + PartitionCache cache(kFsId); + + auto partition = NewPartition(10); + cache.PutIf(partition); + ASSERT_EQ(partition->BaseVersion(), 10); + + // stale base but newer delta still advances the delta + auto stale_base = ShardPartition::New( + mock_processor_, + GenAttrWithMutation(kFsId, kParentIno, 5, {{1, 5}})); + auto result = cache.PutIf(stale_base); + ASSERT_EQ(result.get(), partition.get()); + ASSERT_EQ(partition->BaseVersion(), 10); + ASSERT_EQ(partition->VersionVec().DeltaVersion(1), 5); + ASSERT_EQ(partition->CompleteVersion(), 15); + + // newer base is merged, existing delta is kept + auto newer_base = ShardPartition::New( + mock_processor_, + GenAttrWithMutation(kFsId, kParentIno, 20, {{1, 3}})); + cache.PutIf(newer_base); + ASSERT_EQ(partition->BaseVersion(), 20); + ASSERT_EQ(partition->CompleteVersion(), 25); +} + +TEST_F(ShardPartitionVersionTest, CachePutIfCoversDeltaOpsAndPrunes) { + PartitionCache cache(kFsId); + + auto partition = cache.PutIf(NewPartition(10)); + Dentry dentry( + GenDentry(kFsId, kParentIno, 200, "file1", pb::mds::FileType::FILE)); + partition->Put(dentry, AttrVersion(1, 5)); + ASSERT_EQ(partition->CompleteVersion(), 15); + + auto newer = + ShardPartition::New(mock_processor_, + GenAttrWithMutation(kFsId, kParentIno, 10, {{1, 8}})); + cache.PutIf(newer); + + ASSERT_EQ(partition->CompleteVersion(), 18); + + // the delta op is covered by the merged version and gets pruned + Json::Value value; + partition->Dump(value); + ASSERT_EQ(value["delta_dentry_ops_total"].asUInt64(), 0); +} + class ShardPartitionWithBoundariesTest : public testing::Test { protected: void SetUp() override { @@ -918,7 +1186,7 @@ TEST(ShardPartitionPerfTest, Put400Million) { "avg latency({:.0f}ns)\n", total, elapsed_us / 1e6, qps, elapsed_us * 1000.0 / total); - ASSERT_EQ(partition->DeltaVersion(), total + 1); + ASSERT_EQ(partition->CompleteVersion(), total + 1); } class DirShardConstructFromDentriesTest : public testing::Test { diff --git a/test/unit/mds/storage/test_dummy_storage.cc b/test/unit/mds/storage/test_dummy_storage.cc index 0037dc039..e1782d28b 100644 --- a/test/unit/mds/storage/test_dummy_storage.cc +++ b/test/unit/mds/storage/test_dummy_storage.cc @@ -425,6 +425,85 @@ TEST_F(DummyStorageTest, TxnPutIfAbsentRaceAtCommit) { EXPECT_EQ(v, "v1"); } +TEST_F(DummyStorageTest, TxnWriteConflictOnStaleReadModifyWrite) { + { + auto txn = storage_->NewTxn(); + ASSERT_TRUE(txn->Put("c1", "v0").ok()); + ASSERT_TRUE(txn->Commit().ok()); + } + + // Both txns read the same version, then both write it back. + auto txn1 = storage_->NewTxn(); + auto txn2 = storage_->NewTxn(); + std::string v; + ASSERT_TRUE(txn1->Get("c1", v).ok()); + ASSERT_TRUE(txn2->Get("c1", v).ok()); + ASSERT_TRUE(txn1->Put("c1", "v1").ok()); + ASSERT_TRUE(txn2->Put("c1", "v2").ok()); + + ASSERT_TRUE(txn1->Commit().ok()); + + // txn2's read is stale; its commit must be rejected as retryable. + auto status = txn2->Commit(); + EXPECT_FALSE(status.ok()); + EXPECT_EQ(status.error_code(), pb::error::ESTORE_MAYBE_RETRY); + + ASSERT_TRUE(storage_->Get("c1", v).ok()); + EXPECT_EQ(v, "v1"); + + // A retry that re-reads the latest value commits cleanly. + auto txn3 = storage_->NewTxn(); + ASSERT_TRUE(txn3->Get("c1", v).ok()); + ASSERT_TRUE(txn3->Put("c1", "v3").ok()); + ASSERT_TRUE(txn3->Commit().ok()); + ASSERT_TRUE(storage_->Get("c1", v).ok()); + EXPECT_EQ(v, "v3"); +} + +TEST_F(DummyStorageTest, TxnRecreateAfterDeleteDoesNotConflict) { + // A delete leaves the key absent but keeps its old version in the internal + // version map. A later read-then-write (Lookup(miss) -> create) must still + // commit -- otherwise recreate after delete livelocks. + { + auto txn = storage_->NewTxn(); + ASSERT_TRUE(txn->Put("r1", "v0").ok()); + ASSERT_TRUE(txn->Commit().ok()); + } + { + auto txn = storage_->NewTxn(); + std::string v; + ASSERT_TRUE(txn->Get("r1", v).ok()); + ASSERT_TRUE(txn->Delete("r1").ok()); + ASSERT_TRUE(txn->Commit().ok()); + } + + auto txn = storage_->NewTxn(); + std::string v; + auto status = txn->Get("r1", v); + ASSERT_EQ(status.error_code(), pb::error::ENOT_FOUND); + ASSERT_TRUE(txn->Put("r1", "v1").ok()); + status = txn->Commit(); + EXPECT_TRUE(status.ok()) << status.error_str(); + + ASSERT_TRUE(storage_->Get("r1", v).ok()); + EXPECT_EQ(v, "v1"); +} + +TEST_F(DummyStorageTest, TxnBlindWritesDoNotConflict) { + // Writes without a preceding read stay last-write-wins: the MDS relies on + // this for freshly allocated keys that nobody has read. + auto txn1 = storage_->NewTxn(); + auto txn2 = storage_->NewTxn(); + ASSERT_TRUE(txn1->Put("b1", "v1").ok()); + ASSERT_TRUE(txn2->Put("b1", "v2").ok()); + ASSERT_TRUE(txn1->Commit().ok()); + ASSERT_TRUE(txn2->Commit().ok()); + + std::string v; + ASSERT_TRUE(storage_->Get("b1", v).ok()); + EXPECT_EQ(v, "v2"); +} + TEST_F(DummyStorageTest, ConcurrentTxnsAllCommit) { constexpr int kThreads = 8; constexpr int kPerThread = 50;
RangeIDSizeVersion