-
Notifications
You must be signed in to change notification settings - Fork 1.4k
docs(base): stabilize continuous-period sums #2262
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -53,12 +53,13 @@ lark-cli base +record-list \ | |
| ["关联项目", "intersects", [{"id": "rec_xxx"}]], // 关联项目包含某个 record_id;intersects 表示包含数组中任意一条关联 | ||
| ["备注", "non_empty"], // 格子非空;判断格子为空改用 ["备注", "empty"] | ||
| ["业务日期", "==", "ExactDate(2026-08-07)"], // 具体一天:按 Base 时区匹配 2026-08-07 当天 | ||
| ["发生时间", ">", "ExactDate(2024-01-31 23:59:59.999)"], // 日期不支持 >=;用 > 前一天最后一毫秒表达含当天的下界 | ||
| ["发生时间", "<", "ExactDate(2024-03-01 00:00:00)"] // 2024 年 2 月范围上界:小于 3 月 1 日零点 | ||
| ["发生时间", "<", "ExactDate(2024-03-01)"] // 范围上界:严格小于 3 月 1 日零点,得到 2 月区间的开区间上界;含首日的下界不要用 > 前一天最后一毫秒,改用下文半开区间 [start, next_start) 口径 | ||
|
Collaborator
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 这个设计不能删除,目前不支持 >= 所以必须用 -1ms 来模拟 |
||
| ] | ||
| } | ||
| ``` | ||
|
|
||
| `+record-list` 的 `datetime` 范围条件只使用严格的 `>` / `<`,不要套用通用 tuple 中的 `>=` / `<=`。`ExactDate(...)` 按 Base 时区的日期边界解释,也不能通过把下界前移一天来模拟“包含首日”。需要精确的本地日历半开区间 `[start, next_start)` 时,指标可按日期重组则使用下文 `+data-query` 的日期维度恢复路径;否则按本 SOP 完整导出 NDJSON,再用序列化值中的本地日期筛选。 | ||
|
Collaborator
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 不要加这个 |
||
|
|
||
| 全表分析的常规资源链路是 `+table-list` 确认目标表,并用已有整表 `records_count` 或 NDJSON `has_more` 确认规模;对所有参与分析的表并发执行 `+field-list` 读取所需 schema,再按上述 NDJSON 契约用 `+record-list` 导出记录。已有可信的 `table_id` 时可直接并发读取各表 `+field-list`。`+view-get` 可按需读取,作为用户持久化访问习惯的可选参考;其中的 filter、sort 与字段范围可辅助理解用户常用的查询范围和排序偏好,并结合当前任务确定最终口径。 | ||
|
|
||
| 1. 每次读取使用任务所需的最小投影,并包含 JOIN、解释、回查或写入需要的业务 key。 | ||
|
|
@@ -164,6 +165,10 @@ Agent 上下文曾下载过当前表的 NDJSON 时,按以下规则判断是否 | |
|
|
||
| 例如,`2026-03-20T23:30:00.000-05:00` 与 `2026-03-21T12:30:00.000+08:00` 表示同一时刻;前者若是来源 Base 的值,本地日报归入 3 月 20 日,而时长或排序计算应把它解析为绝对时刻。只构造任务实际需要的日期表示,并在分析引擎中使用具备 datetime 功能的列。 | ||
|
|
||
| 连续日历区间分桶时,从用户请求的范围和粒度生成相邻、不重叠的半开区间 `[start, next_start)`;上界使用下一周期首日,不能把本周期最后一天复用为下一周期起点。 | ||
|
Collaborator
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 冗余了,一个指令没有必要重复那么多次。另外,前一天最后 1ms 的设计被破坏了,不能改动这个 |
||
|
|
||
| 使用 `+data-query` 时,日期 operator、`ExactDate` 以及无法直接表达本地日历边界时的精确恢复路径以 [lark-base-data-query.md](lark-base-data-query.md) 为准;恢复路径不适用时,按上文的谓词探测与 NDJSON 路由完整导出,再用序列化值中的本地日期分桶。无法完整返回或精确表达时,标记证据缺口,不得通过猜测时区或边界补出结果。 | ||
|
|
||
| ## 读取与关系建模 | ||
|
|
||
| 仅在 SOP 已选择 Python 路径后,按实际实现方式只读一份示例: | ||
|
|
@@ -179,6 +184,12 @@ Agent 上下文曾下载过当前表的 NDJSON 时,按以下规则判断是否 | |
|
|
||
| ## 常见分析模式 | ||
|
|
||
| ### 连续分期求和 | ||
|
Collaborator
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 这个不是“常见分析模式”,不要放在这个文档内部 |
||
|
|
||
| - 对按连续日历周期统计的 `sum`,在第一次用于支撑结论的计算前锁定数据源、求和字段、日期字段、完整范围、粒度和过滤条件;主计算、校验和最终答案复用这些口径。 | ||
| - 在把结果视为完整前,确认 `records_count` / `has_more` 或 Cloud 聚合覆盖目标范围,抽查范围边界记录,并将互斥、完整的分桶值之和与同口径全区间结果对账;不一致时先修正范围或计算。 | ||
| - 当用户要求覆盖完整周期的连续分期结果时,最终交付逐桶结果和同口径全区间合计;逐桶值、合计和趋势直接复用同一份已校验的结构化计算结果,提交前逐项核对。用户要求走势时,基于未舍入的分桶值给出相邻周期变化和首尾比较。 | ||
|
Collaborator
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 感觉在过拟合测评题目,移除一下 |
||
|
|
||
| ### 单表简单筛选与统计:jq | ||
|
|
||
| NDJSON 每行是一条 record。单表短筛选、计数和简单聚合可直接用 jq;下面筛选“状态”包含“进行中”的记录,并统计记录数和金额合计: | ||
|
|
||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
这里有问题,目前不支持 isGreaterEqual,所以为了实现包含起始日必须下界前移 1ms,这里不允许前移一天有问题