ai-saas / saas-search

20 Apr, 2026

1 commit

 变更清单

 修复（6 处漂移用例，全部更新到最新实现）
- `tests/test_eval_metrics.py` — 整体重写为新的 4 级 label + 级联公式断言，放弃旧的 `RELEVANCE_EXACT/HIGH/LOW/IRRELEVANT` 和硬编码 ERR 值。
- `tests/test_embedding_service_priority.py` — 补齐 `_TextDispatchTask(user_id=...)` 新必填位。
- `tests/test_embedding_pipeline.py` — cache-hit 路径的 `np.allclose` 改用 `np.asarray(..., dtype=float32)` 避开 object-dtype。
- `tests/test_es_query_builder_text_recall_languages.py` — keywords 次 combined_fields 的期望值对齐现行值（`MSM 60% / boost 0.8`）并重命名。
- `tests/test_product_enrich_partial_mode.py`
  - `test_create_prompt_supports_taxonomy_analysis_kind`：去掉错误假设（fr 不属于任何 taxonomy schema），明确 `(None, None, None)` sentinel 的契约。
  - `test_build_index_content_fields_non_apparel_taxonomy_returns_en_only`：fake 模拟真实 schema 行为（unsupported lang 返回空列表），删除"zh 未被调用"的过时断言。

 清理历史过渡物（per 开发原则：不保留内部双轨）
- 删除 `tests/test_keywords_query.py`（已被 `query/keyword_extractor.py` 生产实现取代的早期原型）。
- `tests/test_facet_api.py` / `tests/test_cnclip_service.py` 移动到 `tests/manual/`，更新 `tests/manual/README.md` 说明分工。
- 重写 `tests/conftest.py`：仅保留 `sys.path` 注入，删除全库无人引用的 `sample_search_config / mock_es_client / test_searcher / temp_config_file` 等 fixture。
- 删除 `tests/test_suggestions.py` 中 13 处残留 `@pytest.mark.unit` 装饰器（模块级 `pytestmark` 已覆盖）。

 新建一致性基础设施
- `pytest.ini`：权威配置源。`testpaths = tests`、`norecursedirs = tests/manual`、`--strict-markers`、登记所有子系统 marker + `regression` marker。
- `tests/ci/test_service_api_contracts.py` + 30 个 `tests/test_*.py` 批量贴上 `pytestmark = [pytest.mark.<subsystem>, pytest.mark.regression]`（AST 安全插入，避开多行 import）。
- `scripts/run_regression_tests.sh` 新建，支持 `SUBSYSTEM=<name>` 选子集。
- `scripts/run_ci_tests.sh` 扩容：由原先的 `tests/ci -q` 改为 `contract` marker + `search ∧ regression` 双阶段。

 文档统一（删除历史双轨）
- 重写 `docs/测试Pipeline说明.md`：删除 `tests/unit/` / `tests/integration/` / `scripts/start_test_environment.sh` 等早已不存在的引用，给出目录约定、marker 表、回归锚点矩阵、覆盖缺口清单、联调脚本用法。
- 删除 `docs/测试回归钩子梳理-2026-04-20.md`（内容已合并进上面一份权威文档，按"一处真相"原则下掉）。
- `docs/DEVELOPER_GUIDE.md §8.2 测试` 改写，指向 pipeline 权威文档。
- `CLAUDE.md` 的 `Testing` 与 `Testing Infrastructure` 两节同步更新。

 最终状态

| 指标 | 结果 |
|------|------|
| 全量 `pytest tests/` | **241 passed** |
| `./scripts/run_ci_tests.sh` | 45 passed |
| `./scripts/run_regression_tests.sh` | 233 passed |
| 子系统子集（示例） | search=45 / rerank=35 / embedding=23 / intent=25 / translation=33 / indexer=17 / suggestion=13 / query=6 / eval=8 / contract=34 |
| 未清零的已知缺口 | 见新版 `测试Pipeline说明.md §4`（function_score / facet / image search / config loader / document_transformer 等 6 条） |

Pipeline 文档里 §4 的覆盖缺口我没有强行补测用例——那属于"新增覆盖"，不是这次清理的范畴；只要后续谁补，把对应 marker 贴上去、从清单里划掉即可。

2026-04-20 12:55:04 +0800

14 Apr, 2026

1 commit

0ba0e0fc 1. rerank漏斗配置优化 ... Browse File »
```
2. +service_enabled_by_config() {
reranker|reranker-fine|translator 如果被关闭，则run.sh all 不启动该服务
```
tangwang
2026-04-14 12:56:25 +0800

27 Mar, 2026

2 commits

5a01af3c 多模态hashkey调整：1. 加入model_name,2.text/url转hash Browse File »

tangwang
2026-03-27 10:36:59 +0800
dc403578 多模态搜索 Browse File »

tangwang
2026-03-27 08:11:35 +0800

23 Mar, 2026

1 commit

4650fcec 日志优化、日志串联（uid rqid） Browse File »

tangwang
2026-03-23 23:45:04 +0800

22 Mar, 2026

1 commit

ef5baa86 混杂语言处理 Browse File »

tangwang
2026-03-22 14:16:39 +0800

20 Mar, 2026

1 commit

b754fd41 图片向量化支持优先级参数 Browse File »

tangwang
2026-03-20 11:59:57 +0800

19 Mar, 2026

1 commit

7214c2e7 mplemented** ... Browse File »

- Text and image embedding are now split into separate
  services/processes, while still keeping a single replica as requested.
The split lives in
[embeddings/server.py](/data/saas-search/embeddings/server.py#L112),
[config/services_config.py](/data/saas-search/config/services_config.py#L68),
[providers/embedding.py](/data/saas-search/providers/embedding.py#L27),
and the start scripts
[scripts/start_embedding_service.sh](/data/saas-search/scripts/start_embedding_service.sh#L36),
[scripts/start_embedding_text_service.sh](/data/saas-search/scripts/start_embedding_text_service.sh),
[scripts/start_embedding_image_service.sh](/data/saas-search/scripts/start_embedding_image_service.sh).
- Independent admission control is in place now: text and image have
  separate inflight limits, and image can be kept much stricter than
text. The request handling, reject path, `/health`, and `/ready` are in
[embeddings/server.py](/data/saas-search/embeddings/server.py#L613),
[embeddings/server.py](/data/saas-search/embeddings/server.py#L786), and
[embeddings/server.py](/data/saas-search/embeddings/server.py#L1028).
- I checked the Redis embedding cache. It did exist, but there was a
  real flaw: cache keys did not distinguish `normalize=true` from
`normalize=false`. I fixed that in
[embeddings/cache_keys.py](/data/saas-search/embeddings/cache_keys.py#L6),
and both text and image now use the same normalize-aware keying. I also
added service-side BF16 cache hits that short-circuit before the model
lane, so repeated requests no longer get throttled behind image
inference.

**What This Means**
- Image pressure no longer blocks text, because they are on different
  ports/processes.
- Repeated text/image requests now return from Redis without consuming
  model capacity.
- Over-capacity requests are rejected quickly instead of sitting
  blocked.
- I did not add a load balancer or multi-replica HA, per your GPU
  constraint. I also did not build Grafana/Prometheus dashboards in this
pass, but `/health` now exposes the metrics needed to wire them.

**Validation**
- Tests passed: `.venv/bin/python -m pytest -q
  tests/test_embedding_pipeline.py
tests/test_embedding_service_limits.py` -> `10 passed`
- Stress test tool updates are in
  [scripts/perf_api_benchmark.py](/data/saas-search/scripts/perf_api_benchmark.py#L155)
- Fresh benchmark on split text service `6105`: 535 requests / 3s, 100%
  success, `174.56 rps`, avg `88.48 ms`
- Fresh benchmark on split image service `6108`: 1213 requests / 3s,
  100% success, `403.32 rps`, avg `9.64 ms`
- Live health after the run showed cache hits and non-zero cache-hit
  latency accounting:
  - text `avg_latency_ms=4.251`
  - image `avg_latency_ms=1.462`

2026-03-19 13:21:01 +0800

17 Mar, 2026

2 commits

4a37d233 1. embedding cache float32 -> bf16 ... Browse File »

2. 抽象出可复用的 embedding Redis 缓存类（图文共用）

详细：
1. embedding 缓存改为 BF16 存 Redis（读回恢复 FP32）
关键行为（按你给的流程落地）
写入前：FP32 embedding →（normalize_embeddings=True 时）L2 normalize →
转 BF16 → bytes（2字节/维，大端） → redis.setex
读取后：redis.get bytes → BF16 → 恢复 FP32（np.float32 向量）
变更点
新增 embeddings/bf16.py
提供 float32_to_bf16 / bf16_to_float32
encode_embedding_for_redis()：FP32 → BF16 → bytes
decode_embedding_from_redis()：bytes → BF16 → FP32
l2_normalize_fp32()：按需归一化
修改 embeddings/text_encoder.py
Redis value 从 pickle.dumps(np.ndarray) 改为 BF16 bytes
缓存 key 改为包含 normalize 标记：{prefix}:{n0|n1}:{query}（避免
normalize 开关不同却共用缓存）
修改 tests/test_embedding_pipeline.py
cache hit 用例改为写入 BF16 bytes，并使用新
key：embedding:n1:cached-text
修改 docs/缓存与Redis使用说明.md
embedding 缓存的 Key/Value 格式更新为 BF16 bytes + n0/n1
修改 scripts/redis/redis_cache_health_check.py
embedding pattern 不再硬编码 embedding:*，改为读取
REDIS_CONFIG["embedding_cache_prefix"]
value 预览从 pickle 解码改为 BF16 解码后展示 dim/bytes/dtype
自检
在激活环境后跑过 BF16 编解码往返 sanity check：bytes
长度、维度恢复正常；归一化向量读回后范数接近 1（会有 BF16 量化误差）。

2. 抽象出可复用的 embedding Redis 缓存类（图文共用）
新增
embeddings/redis_embedding_cache.py：RedisEmbeddingCache
统一 Redis 初始化（读 REDIS_CONFIG）
统一 BF16 bytes 编解码（复用 embeddings/bf16.py）
统一过期策略：写入 setex(expire_time)，命中读取后 expire(expire_time)
滑动过期刷新 TTL
统一异常/坏数据处理：解码失败或向量非 1D/为空/含 NaN/Inf 会删除该 key
并当作 miss
已接入复用
文本 embeddings/text_encoder.py
用 self.cache = RedisEmbeddingCache(key_prefix=..., namespace="")
key 仍是：{prefix}:{query}
图片 embeddings/image_encoder.py
用 self.cache = RedisEmbeddingCache(key_prefix=..., namespace="image")
key 仍是：{prefix}:image:{url_or_path}

2026-03-17 15:06:51 +0800

3d588bef embeddings Browse File »

tangwang
2026-03-17 13:53:50 +0800

13 Mar, 2026

2 commits

d4cadc13 翻译重构 Browse File »

tangwang
2026-03-13 20:28:08 +0800
77ab67ad 更新测试用例 Browse File »

tangwang
2026-03-13 12:39:40 +0800

10 Mar, 2026

1 commit

24e92141 delete enable_multilang_search Browse File »

tangwang
2026-03-10 13:12:56 +0800

09 Mar, 2026

2 commits

ed948666 tidy Browse File »

tangwang
2026-03-09 17:04:00 +0800
950a640e embeddings Browse File »

tangwang
2026-03-09 15:59:14 +0800