
数据分析机器学习数据科学【免费下载链接】KatsKats, a kit to analyze time series data, a lightweight, easy-to-use, generalizable, and extendable framework to perform time series analysis, from understanding the key statistics and characteristics, detecting change points and anomalies, to forecasting future trends.项目地址https://gitcode.com/gh_mirrors/ka/Kats点击查看免费下载本文以 Kats 开源时序分析框架中的 prophet_detector 模块 为对象系统讲解如何将 Prophet 时序预测模型封装为异常检测模型ProphetDetectorModel与趋势检测模型ProphetTrendDetectorModel。读完本文你将掌握该检测器的完整参数体系、评分策略偏差率与 Z 分数、季节性/节假日建模、训练区间排除、序列化等核心能力并能直接落地到真实业务数据上。该模块的官方 API 文档入口为 kats.detectors.prophet_detector.rst本文内容即围绕该文档所对应的模块实现展开。一、模块定位把预测区间变成异常分数ProphetDetectorModel的基本思想非常直观先用 Prophet 库对历史数据建模预测下一个时间点的取值区间再将实际观测值与预测区间进行比较偏差越大异常分数越高。模块源码开头的文档字符串对此有明确描述A Detector Model that does anomaly detection, by first using the Prophet library to forecast the interval for the next point, and comparing this to the actually observed data point.从源码结构看ProphetDetectorModel继承自 DetectorModel 基类因此遵循 Kats 检测器统一的生命周期协议fit(data, historical_data)训练模型predict(data, historical_data)仅预测不改状态fit_predict(data, historical_data)先训练再预测serialize()将模型序列化为 JSON 字节串便于保存与恢复。由于ProphetDetectorModel通过 DetectorModelRegistry 自动注册还可以用类名从注册表中直接实例化。值得注意的是该模块是 Kats 的可选依赖模块在 kats/detectors/init.py 中prophet_detector的导入被try/except ImportError包裹未安装 Prophet 时仅给出 warning。因此使用前需先安装prophet见 requirements.txt。二、核心类ProphetDetectorModel 全参数详解ProphetDetectorModel的构造函数位于 prophet_detector.py包含以下核心参数参数默认值说明serialized_modelNone此前序列化模型产生的 JSON 字节串传入后直接反序列化复用模型score_funcdeviation_from_predicted_val评分函数可选deviation_from_predicted_val或z_score支持字符串与枚举两种写法scoring_confidence_interval0.8Prophet 的interval_width部分评分策略用它计算异常分数remove_outliersFalse训练前是否先剔除离群点outlier_threshold0.99离群点剔除时的置信区间阈值uncertainty_samples50Prophet 计算不确定性所需的采样次数outlier_removal_uncertainty_samples40离群点剔除阶段使用的采样次数vectorizeFalse是否使用向量化方法生成预测区间可显著提升性能use_legacy_z_scoreTrue使用旧版 Z 分数计算方式默认值下与新版一致seasonalitiesNone季节性配置见下节country_holidaysNone国家代码如US自动加入该国法定节假日holidays_listNone自定义节假日日期列表或{ds: [...], holiday: [...]}字典holiday_multiplierNone节假日期间异常分数的乘数exclude_training_rangesNone从训练数据中排除的时间范围列表如[[1672552800, 1672567200]]saturation_rangeNone饱和度区间[min, max]指定后使用 logistic 增长模型growth_typelinear趋势增长类型可选linear、flat、logistic2.1 构造函数中的内部逻辑从源码实现中可以观察到几个值得注意的行为score_func 字符串兼容为兼容参数调优param tuning框架score_func接受字符串形式内部通过STR_TO_SCORE_FUNC字典映射prophet_detector.py无法识别的字符串会回退到默认评分函数。uncertainty_samples 的动态归零只有当使用z_score评分时uncertainty_samples才保持用户设定值使用默认评分偏差率时会被强制设为0从而跳过置信带计算以提升运行性能prophet_detector.py。saturation_range 与 growth_type 的联动若传入saturation_range而growth_type不是logistic会打印 warning 并自动把growth_type改为logisticprophet_detector.pysaturation_range必须是有序的两个数值否则抛出ValueError见 test_saturation_range_validation。growth_type 校验非法取值如quadratic会抛出ParameterError对应测试见 test_unsupported_growth_type_throws_error。三、训练、预测与序列化生命周期 API3.1 fit训练 Prophet 模型from kats.consts import TimeSeriesData from kats.detectors.prophet_detector import ProphetDetectorModel model ProphetDetectorModel() model.fit(ts[:90]) # 用前 90 个点训练fit的内部流程prophet_detector.py依次为合并data与historical_data若有作为训练全集若指定了exclude_training_ranges先剔除相应区间全部被剔除则抛DataError将TimeSeriesData转换为 Prophet 需要的{ds, y}DataFrame转换函数见 timeseries_to_prophet_df仅支持单变量数据若指定saturation_range为数据追加floor、cap列若remove_outliersTrue先做一轮离群点剔除见第五节根据数据时间范围自动处理季节性配置组装并fitProphet 模型。训练过程中Kats 默认通过SilentStdoutStderr上下文管理器抑制 Prophet 拟合时的大量日志输出prophet_detector.py如需保留日志可设置环境变量NOT_SUPPRESS_PROPHET_FIT_LOGS1。3.2 predict / fit_predict产出 AnomalyResponseres model.predict(ts[90:]) # 仅预测 res model.fit_predict(datats[90:], historical_datats[:90]) # 先训练再预测fit_predict要求必须提供historical_data否则抛出DataInsufficientErrorprophet_detector.py。predict在模型未训练时调用会抛出InternalError。predict返回 AnomalyResponse 对象包含五个字段scores异常分数时间序列与输入数据等长confidence_band上下置信带ConfidenceBand(lower, upper)predicted_tsProphet 预测值anomaly_magnitude_ts异常幅度当前实现为全零占位stat_sig_ts统计显著性当前实现为全零占位。测试用例 test_no_anomaly_prediction_length 验证了预测分数与输入数据长度一致test_find_anomaly_when_anomaly_present 则验证模型在存在异常时能给出超过阈值±0.3的分数而干净数据不会触发误报。3.3 serialize模型持久化与复用serialized model.serialize() # bytes model2 ProphetDetectorModel(serialized_modelserialized) # 恢复serialize基于 Prophet 官方的model_to_json/model_from_json实现prophet_detector.py。由于 Prophet 序列化在带时区数据上曾出现过兼容性问题模块提供了容错加载函数 load_model_from_json首次反序列化失败时会用正则把时间字符串中的时区标记Z剥离后再尝试。测试 test_serialized_prophet_version_key 确认序列化结果中包含__prophet_version版本键。四、评分策略偏差率 vs Z 分数评分函数通过ProphetScoreFunction枚举与SCORE_FUNC_DICT注册prophet_detector.pydeviation_from_predicted_val默认(实际值 - 预测值) / |预测值|即相对预测值的偏差比例。实现见 deviation_from_predicted_val。z_score(实际值 - 预测值) / 缩放标准差其中标准差由 Prophet 置信带宽度推导而来实现见 z_score。当标准差接近 0 时会被钳制到MIN_STDEV 1e-9避免除零错误。两个常量的设计值得一提Z_SCORE_CI_THRESHOLD_SCALE_CONST norm.ppf(0.8 / 2 0.5) / 0.8 Z_SCORE_SCALE_CONST (50 ** 0.5) * Z_SCORE_CI_THRESHOLD_SCALE_CONST / 2源码注释说明旧实现误将置信带宽度当作 Z 统计量直接使用这两个缩放常量保证在默认参数下修正后的 Z 分数与原分数一致而在自定义置信区间时表现为真正的 Z 分数。默认的uncertainty_samples50与scoring_confidence_interval0.8正是与之配套的默认值。选择z_score评分时confidence_band返回真实的上下界而使用默认偏差率评分时置信带会被设置为与predicted_ts相同prophet_detector.py。这一行为在测试 test_default_score_func 与 test_score_func_parameter_as_z_score 中被精确断言。Z 分数策略的额外特性均有对应测试验证异常幅度越大Z 分数越高test_z_score_proportional_to_anomaly_magnitude平坦信号与零噪声信号不会出现除零错误test_flat_signal、test_zero_noise_signal。五、季节性建模DAY / WEEK / YEAR / WEEKENDSeasonalityTypes枚举定义了四种季节性类型prophet_detector.pyclass SeasonalityTypes(Enum): DAY 0 WEEK 1 YEAR 2 WEEKEND 3seasonalities参数支持三种传法由 seasonalities_to_dict 归一化处理# 1. 单个枚举开启该季节性 ProphetDetectorModel(seasonalitiesSeasonalityTypes.WEEKEND) # 2. 枚举/字符串列表逐个开启 ProphetDetectorModel(seasonalities[SeasonalityTypes.WEEKEND, DAY]) # 3. 字典显式控制 True / False / auto ProphetDetectorModel(seasonalities{SeasonalityTypes.WEEKEND: True, SeasonalityTypes.DAY: auto})5.1 季节性自动处理规则seasonalities_processingprophet_detector.py在每次fit时根据当前时间序列的实际时间跨度动态决定各季节性的最终取值DAY、YEAR 默认auto交由 Prophet 自动判断WEEK 默认auto但若 WEEKEND 被显式开启则 WEEK 自动设为False避免与周末模型重叠WEEKEND 默认False若设为auto当序列总时长不足两周或最小采样间隔大于等于一周时自动降级为False除True/False/auto之外的取值会抛出ParameterError。5.2 周末季节性WEEKENDWEEKEND 是 Kats 在 Prophet 之上的一个增强通过 prophet_weekend_masks 生成weekend_mask与workday_mask两个条件傅里叶季节性周期 7、傅里叶阶数 3使模型能分别刻画工作日与周末的不同模式。测试 test_weekend_seasonality_noise_signal 证明对于工作日/周末模式显著不同的序列开启 WEEKEND 季节性可以显著降低预测 MAE。六、节假日建模国家节假日、自定义日期与分数加权ProphetDetectorModel提供三层节假日能力1. 国家法定节假日model ProphetDetectorModel(country_holidaysUS)内部通过 Prophet 的make_holidays_df为训练数据覆盖的所有年份生成该国节假日并调用model.add_country_holidays()注入。2. 自定义节假日日期# 简单形式日期字符串列表统一命名为 user_provided_holiday ProphetDetectorModel(holidays_list[2022-01-01, 2022-03-31]) # 复杂形式支持不同节假日模式的字典 ProphetDetectorModel(holidays_list{ ds: [2022-01-01, 2022-03-31], holiday: [playoff, superbowl], })传入的列表会在fit时被转换为{ds: [...], holiday: [...]}形式见 prophet_detector.py。3. 节假日分数加权holiday_multiplier节假日往往对应业务高峰holiday_multiplier允许对节假日当天的异常分数做乘法加权例如置零以豁免节假日异常model ProphetDetectorModel(holiday_multiplier0.0)predict阶段会通过 get_holiday_dates 汇总自定义与国家级节假日按日粒度匹配并乘以乘数prophet_detector.py。测试 test_heteroskedastic_noise_signal_with_specific_holidays_mulitplier 验证了乘数为 0 时节假日当天分数归零、次日恢复正常的行为。七、训练数据预处理离群点剔除与区间排除7.1 离群点自动剔除remove_outliersmodel ProphetDetectorModel(remove_outliersTrue, outlier_threshold0.99)_remove_outliersprophet_detector.py的实现思路是两阶段建模先用一个临时 Prophet 模型interval_widthoutlier_threshold默认 0.99对原始数据做拟合预测凡落在预测置信带之外的观测点即视为离群点并从训练集中剔除。该阶段使用独立的采样参数outlier_removal_uncertainty_samples默认 40。测试验证了两点效果剔除离群点后训练的模型其预测 RMSE 不高于未剔除的模型test_outlier_removal_efficacyoutlier_threshold有moderate0.99与aggressive0.8两档可选阈值越低剔除越激进test_outlier_removal_threshold对不规整采样间隔的序列同样安全不会抛ValueErrortest_irregular_data_intervals。7.2 训练区间排除exclude_training_ranges当历史数据中存在已知的异常扰动如人为操作的等级跳变时可以显式排除相应时间段避免污染模型model ProphetDetectorModel( exclude_training_ranges[[1672552800, 1672567200]], # 支持 Unix 秒 ) # 也支持 pandas Timestamp 形式_exclude_rangesprophet_detector.py会将整数时间戳按units、originunix转换为 pandas 时间戳并保留原始序列的时区设置再通过TimeSeriesData.exclude()剔除。测试 test_exclude_ts_range_from_model_training 证明用该参数剔除数据与手工先剔除再训练的结果完全一致误差小于 1e-5。八、增长模型与饱和度linear / flat / logisticgrowth_type与saturation_range共同控制趋势建模方式GrowthType 枚举linear默认线性趋势与 Prophet 默认行为一致保持向后兼容flat零斜率趋势模型只拟合均值水平logistic逻辑斯蒂增长需要同时指定saturation_range[min, max]作为容量上界与下界对应 Prophet 的cap/floor。model ProphetDetectorModel(growth_typelogistic, saturation_range[0.0, 100.0])相关测试提供了有力的行为证据用趋势数据训练时flat模型斜率k0linear模型斜率k≈1且 linear 预测序列尾值大于头值test_flat_vs_linear_growth_when_trained_on_trended_data线性模型预测可能大幅越出边界如低于saturation_min或高于saturation_max而 logistic 饱和度模型会把预测约束在边界附近test_saturation_range_enforcement序列化结果中的growth字段会如实记录所选类型test_growth_types_linear_and_flat_are_passed_to_model。九、趋势检测ProphetTrendDetectorModel同模块还提供了基于 Prophet 的趋势变化点检测模型ProphetTrendDetectorModelprophet_detector.py其构造函数参数为参数默认值说明serialized_modelNone已序列化模型changepoint_range1.0历史数据中用于估计趋势变化点的比例weekly_seasonalityauto周季节性配置changepoint_prior_scale0.01变化点先验尺度控制趋势变化的灵活度与异常检测模型不同该模型只实现fit_predictfit与predict直接抛出NotImplementedError。它的工作原理是拟合 Prophet 后取出模型估计的变化点model.changepoints将每个变化点处斜率增量model.params[delta]各采样均值的绝对值作为该时间点的分数其余时间点分数为 0最终返回AnomalyResponse其中scores即趋势强度序列。测试 test_response_shape_for_single_series 与 test_pmm_use_case 验证了返回形状与历史数据/当前数据的拼接逻辑。十、向量化预测与相关配套实现vectorize参数控制预测区间生成方式其底层实现在 kats/models/prophet.py 中与 Kats 对 Prophet 的预测增强一脉相承非向量化逐条调用 Prophet 原生sample_model采样向量化通过 sample_posterior_predictive 与 sample_model_vectorized 批量生成后验样本其中趋势不确定性由 predict_uncertainty 汇总为百分位区间。测试 test_fit_predict 验证了两种模式产出的分数完全一致同时覆盖了三种预测场景训练/测试时间连续、测试起点晚于训练终点存在时间间隔、测试终点早于训练终点回看历史窗口。此外Kats 还为 Prophet 提供了常规预测模型封装 kats/models/prophet.py 中的 ProphetModel 与ProphetParams参数类适合纯预测场景本文讨论的检测器则专注于异常/趋势检测任务二者可互补使用。十一、完整实战示例下面综合以上内容给出一个覆盖主要能力的端到端示例数据加载方式参考 kats/data/utils.py 与测试用例import pandas as pd from kats.consts import TimeSeriesData from kats.data.utils import load_air_passengers from kats.detectors.prophet_detector import ( ProphetDetectorModel, ProphetScoreFunction, ProphetTrendDetectorModel, SeasonalityTypes, ) # 1. 加载单变量时序数据 data load_air_passengers(return_tsFalse) # {ds, y} 形式 ts TimeSeriesData(data) # 2. 训练-测试切分 train_ts, test_ts ts[:120], ts[120:] # 3. 异常检测Z 分数评分 周末季节性 美国节假日 离群点剔除 model ProphetDetectorModel( score_funcProphetScoreFunction.z_score, # 也支持字符串 z_score scoring_confidence_interval0.8, seasonalities{SeasonalityTypes.WEEKEND: True}, country_holidaysUS, remove_outliersTrue, outlier_threshold0.99, vectorizeTrue, ) response model.fit_predict(datatest_ts, historical_datatrain_ts) # 4. 结果解析 scores response.scores # 异常分数时序 pred response.predicted_ts # 预测值 upper response.confidence_band.upper # 上置信带 lower response.confidence_band.lower # 下置信带 # 5. 模型持久化与恢复 blob model.serialize() restored ProphetDetectorModel(serialized_modelblob) # 6. 趋势变化点检测 trend_detector ProphetTrendDetectorModel( changepoint_range0.8, changepoint_prior_scale0.01, ) trend_resp trend_detector.fit_predict(datats, historical_dataNone) trend_scores trend_resp.scores # 变化点处为趋势强度其余为 0十二、小结与使用建议围绕 kats.detectors.prophet_detector.rst 所对应的模块实现本文完整覆盖了 Kats 中 Prophet 检测器的全部能力异常检测ProphetDetectorModel提供偏差率/Z 分数两种评分、WEEKEND 增强季节性、国家/自定义节假日、训练离群点剔除、区间排除、logistic 饱和度与模型序列化趋势检测ProphetTrendDetectorModel基于 Prophet 变化点估计输出趋势强度序列性能与工程vectorize向量化预测、日志抑制NOT_SUPPRESS_PROPHET_FIT_LOGS、DetectorModelRegistry自动注册。实战选型建议当序列带有明显的周末/节假日效应时优先开启 WEEKEND 季节性与country_holidays当训练数据包含已知扰动区间时使用exclude_training_ranges而非手工清洗对噪声较大的业务数据建议开启remove_outliers需要跨进程保存检测器状态时务必使用serialize()/serialized_model组合。更多源码细节可继续阅读 prophet_detector.py 及其配套测试 test_prophet_detector.py。赞分享数据分析机器学习数据科学【免费下载链接】KatsKats, a kit to analyze time series data, a lightweight, easy-to-use, generalizable, and extendable framework to perform time series analysis, from understanding the key statistics and characteristics, detecting change points and anomalies, to forecasting future trends.项目地址https://gitcode.com/gh_mirrors/ka/Kats点击查看免费下载相关推荐LeetCode 694 Number of Distinct Islands 全解坐标归一化与三种形状哈希去重方案LeetCode 694 Number of Distinct Islands 全解坐标归一化与三种形状哈希去重方案 本篇文章基于当前仓库的题解文档 numb数据分析机器学习数据科学Kats异常检测完全手册从CUSUM到Prophet的10种方法Kats异常检测完全手册从CUSUM到Prophet的10种方法 时间序列异常检测是数据分析中的关键任务而Kats作为一个轻量级、易用且功能强大的时间序列分数据分析机器学习数据科学Kats时间序列预测模型对比Prophet vs ARIMA vs LSTMKats时间序列预测模型对比Prophet vs ARIMA vs LSTM 在当今数据驱动的时代时间序列预测已成为企业决策和业务分析的关键工具。Kats作数据分析机器学习数据科学上一篇7 天 Token 有效怎么让用户不再次输密码Bangumi 的 OAuth 登录实现下一篇PDF-Extract-Kit 扩展开发指南基于注册表机制新增 YOLO 布局检测任务与模型创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考