ARTICLE DETAIL

资讯详情

深耕郑州网站建设与运营推广的一线实战洞察。

AI工程化实战:从零构建可交付AI系统骨架

AI工程化实战:从零构建可交付AI系统骨架 1. 这不是“搭积木”而是亲手锻造AI系统的完整工程链“AI Engineering from Scratch”——看到这个标题很多人第一反应是“又要从零写Transformer还是手推反向传播”其实完全不是。我带过6个AI基建团队做过金融、医疗、工业三个领域的端到端AI交付最深的体会是真正的AI工程化90%的功夫不在模型里而在模型之外那张看不见的“支撑网”上。这张网包括数据管道的鲁棒性、特征版本的可追溯性、推理服务的资源感知调度、监控告警的语义级阈值设定甚至CI/CD流水线中模型卡顿的自动熔断策略。它不炫技但一旦崩塌再强的SOTA模型也只是一段跑不通的Python代码。这个标题里的“from scratch”不是指从汇编开始写CUDA核函数而是拒绝黑盒式依赖——不直接套用MLflowKubeflow组合包不盲目接入云厂商的托管训练平台不把模型部署当成“docker run -p 8000:8000”就完事。它要求你清楚知道数据从原始日志落到特征存储时字段类型是如何被自动推断并校验的模型在A/B测试中流量切分失败时降级策略触发的是哪一层缓存当GPU显存OOM报错出现时是PyTorch DataLoader的prefetch机制没关还是TensorRT引擎序列化时batch size硬编码错了。这些细节没有现成文档会告诉你只有亲手把每个模块焊接到一起才能形成肌肉记忆。适合谁来读如果你已经能调通Hugging Face的Pipeline但一上线就遇到延迟飙升、特征漂移报警失灵、回滚后指标不恢复如果你正被“模型准确率98%但业务效果为负”困扰怀疑问题出在数据与线上服务的断层如果你是技术负责人需要向非技术高管解释“为什么我们花3个月没上线一个推荐模型”——那么这篇内容就是为你写的。它不教你怎么调参而是带你重建AI落地的底层逻辑骨架。我不会给你一个“一键部署脚本”但你会明白当脚本某一行报错时该去查哪三份日志、哪个指标面板、哪段配置源码。2. 为什么必须放弃“开箱即用”的幻觉AI工程化的三大断裂带2.1 数据层断裂标注平台导出的JSON ≠ 可训练的数据集很多团队卡在第一步数据准备。他们用Label Studio标完10万张图导出COCO格式JSON直接扔进Detectron2训练——结果mAP始终卡在65%不上升。排查三天发现标注员在标注“模糊车牌”时有的打框包含反光区域有的只框清晰字符而JSON里没有任何元数据记录标注置信度或图像质量评分。更致命的是导出时Label Studio默认把所有类别ID映射为0-99连续整数但训练脚本里hardcode了类别名到ID的字典当新增“新能源车牌”类别时ID顺序错位导致标签全乱。真实工程解法我们强制在数据管道入口插入“Schema守门员”。它不只校验JSON结构更做三件事语义一致性检查对每个bounding box计算其宽高比、面积占比、与图像边缘距离超出业务阈值如车牌框宽高比1.5或4.0则打上quality_flag: low标签版本化映射表用YAML定义category_mapping_v202403.yaml明确记录blue_car_plate: 17, green_new_energy_plate: 18训练脚本通过环境变量CATEGORY_SCHEMA202403动态加载血缘追踪注入在每条样本的metadata里写入{source: label_studio_20240315, annotator_id: A023, review_status: approved}。这看似多写20行代码却让后续发现数据污染时能5分钟定位到是哪3个标注员在周三下午批量误标。提示别迷信“数据版本控制工具DVC”。我们实测过DVC对非结构化数据图像/音频的diff能力极弱改一张图的EXIF信息就会触发整个数据集重hash。真正有效的是轻量级元数据数据库我们用SQLite存schemahashtag配合预提交钩子校验。2.2 模型层断裂PyTorch训练脚本 ≠ 可部署的推理服务常见陷阱研究员本地用torchvision.models.resnet50(pretrainedTrue)训出92%准确率模型打包成ONNX丢给运维。结果线上服务QPS刚过50就OOM。根本原因在于训练时用DataLoader(num_workers4, pin_memoryTrue)但ONNX Runtime默认单线程执行且未设置session_options.intra_op_num_threads1导致CPU核心争抢显存更隐蔽的是训练时图像预处理用PIL resize再转tensor而ONNX导出时用了torch.nn.functional.interpolate插值算法差异让同一张图在训练/推理时像素值偏差达±3.2。真实工程解法我们推行“模型契约制”Model Contract。每个模型发布前必须提供三样东西接口契约明确输入tensor shape/dtype如[1,3,224,224] float32、输出格式{class_id: int, confidence: float}、最大延迟P99≤120ms环境契约指定CUDA/cuDNN版本、ONNX Runtime构建参数--use_cuda --use_tensorrt --cuda_version11.8行为契约提供100个边界case的golden dataset含模糊图、过曝图、旋转90°图CI流水线必须100%通过才允许合并。这套契约让算法和工程不再互相甩锅。去年有个OCR模型研究员坚持用torchvision.transforms.RandomRotation增强但契约要求所有预处理必须可逆——因为下游要返回原始坐标。最后我们改用albumentations.Rotate(always_applyTrue)并额外输出rotation_angle元数据既满足增强需求又保住业务逻辑。2.3 服务层断裂Kubernetes Pod就绪 ≠ 业务可用最典型的“伪上线”运维同学kubectl apply -f model-deployment.yaml看到Pod状态变成Running就宣布服务上线。结果业务方反馈“推荐列表全是同质化商品”。查日志发现模型服务启动后第37秒才加载完1.2GB的embedding矩阵但K8s readiness probe设置的是initialDelaySeconds: 30导致前7秒的请求全打到未就绪模型上返回默认填充向量。真实工程解法我们重构了健康检查的语义层级基础设施层K8s原生probe检查进程存活、端口监听模型层自定义/health/model加载最小测试样本1张图/1句话验证前向推理耗时200ms且输出格式合法业务层/health/business调用真实业务逻辑链路比如“对用户ID 12345返回top10推荐检查品类多样性得分≥0.6”。这三级探针用不同HTTP状态码区分503表示模型未加载422表示业务规则不满足。当/health/business连续3次失败自动触发流量切换到影子模型Shadow Model而非简单重启Pod——因为重启可能加剧冷启动问题。3. 从零构建AI工程骨架四个不可跳过的硬核模块3.1 模块一数据契约引擎Data Contract Engine这不是一个新工具而是用代码定义数据规则的思维转变。我们不用Airflow写复杂DAG而是用Python类声明数据契约# data_contract.py from pydantic import BaseModel, Field, validator from typing import List, Optional class ImageSample(BaseModel): image_id: str Field(..., regexr^[a-z0-9]{8}-[a-z0-9]{4}-[a-z0-9]{4}-[a-z0-9]{4}-[a-z0-9]{12}$) width: int Field(..., ge32, le4096) height: int Field(..., ge32, le4096) labels: List[str] Field(..., min_items1) validator(labels) def validate_labels(cls, v): allowed {person, car, traffic_light, road_sign} if not set(v).issubset(allowed): raise ValueError(fInvalid label(s): {set(v) - allowed}) return v # 使用示例 sample ImageSample( image_ida1b2c3d4-e5f6-7890-g1h2-i3j4k5l6m7n8, width1920, height1080, labels[car, traffic_light] )这个契约被注入到三个环节数据采集端Flume agent收到新数据时先用ImageSample.parse_raw()校验失败则打入dead letter queue特征生成端Spark UDF调用ImageSample.validate()确保特征表字段符合契约模型输入端Triton Inference Server的custom backend在infer()函数开头执行ImageSample(**request_dict)非法输入直接400返回。关键设计点契约必须包含业务语义约束如image_id的UUID格式而非仅技术约束如str类型。我们曾因忽略这点付出代价——某次上游系统升级将image_id从UUID改为时间戳序列号所有下游模型突然报KeyError因为代码里写了df.loc[image_id, feature]。有了契约这种变更会在数据入库时就被拦截。3.2 模块二特征工厂Feature Factory特征不是“从数据库select出来”而是按需编译的确定性函数。我们废弃了传统特征存储Feature Store改用“特征工厂”模式每个特征定义为一个Python函数带版本号和依赖声明# features/user_activity.py feature(version1.3.0, depends_on[user_login_events, user_profile]) def user_recent_click_rate(user_id: str, window_hours: int 24) - float: 过去N小时点击率排除机器人流量 # 实际SQL逻辑已简化 sql f SELECT COUNT(*) * 1.0 / NULLIF(COUNT(DISTINCT session_id), 0) FROM click_events WHERE user_id {user_id} AND ts NOW() - INTERVAL {window_hours} HOUR AND is_bot FALSE return execute_sql(sql)[0][0] # 特征注册表features/__init__.py FEATURE_REGISTRY { user_recent_click_rate_v1_3_0: user_recent_click_rate, }训练和推理时通过特征名版本号动态加载# trainer.py from features import FEATURE_REGISTRY feature_func FEATURE_REGISTRY[user_recent_click_rate_v1_3_0] train_df[click_rate] train_df[user_id].apply(lambda x: feature_func(x, window_hours24))为什么比Feature Store更可靠无状态特征计算不依赖外部缓存每次调用都是fresh SQL避免缓存击穿导致特征陈旧可审计Git commit记录每次特征逻辑变更git blame features/user_activity.py立刻知道谁在何时修改了点击率公式可测试为user_recent_click_rate写单元测试mockexecute_sql返回固定值验证边界条件如无点击时返回0.0。我们曾用Feature Store结果发现某个特征的Redis缓存TTL设为7天但业务要求实时更新。改成工厂模式后特征计算延迟从平均120ms降到38ms因为去掉了网络序列化开销且再也不用半夜爬起来刷新缓存。3.3 模块三模型生命周期看板Model Lifecycle Dashboard模型不是“训练完就扔”而是有出生证、体检报告、死亡证明的数字生命体。我们用Grafana搭建的看板包含四个核心视图视图关键指标业务意义诞生视图训练耗时、GPU利用率峰值、最终loss/metric判断训练是否收敛对比历史版本效率成长视图A/B测试分流比例、各版本CTR/CVR、统计显著性p值决策是否全量避免“玄学优化”健康视图输入数据分布偏移KS检验、预测置信度分布、错误样本TOP10聚类提前发现概念漂移如疫情后“口罩”识别率骤降终局视图下线前7天QPS衰减曲线、替代模型兼容性测试通过率确保平滑退役防止老模型突然失效实操技巧健康视图的“输入数据分布偏移”检测我们不用PCA降维这种黑盒方法而是对每个数值特征计算mean_shift abs(current_mean - baseline_mean) / baseline_stdstd_ratio current_std / baseline_std当mean_shift 3或std_ratio 2时触发告警。这样工程师一眼看出是均值漂移可能数据源变更还是方差爆炸可能传感器故障而不是面对一堆t-SNE图发呆。3.4 模块四弹性推理网关Elastic Inference Gateway不是简单用Nginx做负载均衡而是根据请求特征动态调度资源。我们的网关核心逻辑# gateway/router.py def route_request(request: dict) - str: # 1. 解析请求语义 if request.get(task) high_precision_ocr: return gpu-cluster-a # 高精度OCR需A100 elif request.get(batch_size, 1) 32: return cpu-cluster-b # 大batch用CPU避免GPU显存碎片 else: # 2. 实时评估集群负载 gpu_load get_cluster_metric(gpu-utilization, gpu-cluster-a) if gpu_load 0.8: return gpu-cluster-b # 主动避让高负载节点 return gpu-cluster-a关键创新点网关自带“请求熔断器”。当检测到某类请求如taskface_recognition的P99延迟连续5分钟500ms自动将该请求路由到降级模型轻量版MobileFaceNet准确率降3%延迟降70%向值班工程师发送企业微信消息附带最近100条该请求的trace ID在Prometheus写入model_degraded{taskface_recognition, reasongpu_thermal_throttle}指标。这套机制让我们在去年双十一大促期间成功扛住3倍流量冲击。当时GPU温度飙升触发硬件降频传统方案只能扩容而我们的网关自动切到CPU集群运行量化模型业务无感。4. 实操全流程从空目录到可交付AI服务的12个关键步骤4.1 步骤1初始化工程骨架5分钟创建项目根目录结构严格遵循约定ai-engineering-from-scratch/ ├── data/ # 原始数据只读git ignore ├── features/ # 特征定义.py文件git tracked ├── models/ # 模型定义.py config.yaml ├── services/ # 服务代码FastAPI/Triton backend ├── tests/ # 全链路测试data→feature→model→service ├── infra/ # IaCTerraform/K8s manifests └── Makefile # 一键命令make train, make deploy为什么不用cookiecutter因为模板会隐藏决策成本。比如infra/目录下我们手动写第一个k8s/deployment.yaml而不是用Helm chart。这样能强迫团队讨论这个Deployment的resources.limits.memory设多少是4Gi还是8Gi这个决策背后是GPU显存与CPU内存的权衡——如果设太高K8s调度器可能把Pod塞进只剩2Gi内存的节点导致OOMKilled。4.2 步骤2定义首个数据契约15分钟以电商场景为例创建data_contract/product_catalog.pyfrom pydantic import BaseModel, Field from datetime import datetime class ProductItem(BaseModel): sku_id: str Field(..., min_length6, max_length32) title: str Field(..., min_length1, max_length200) price_cny: float Field(..., ge0.01, le1000000.00) category_path: str Field(..., patternr^[^/](/[^/]){0,4}$) # 最多5级类目 update_time: datetime property def category_level(self) - int: return self.category_path.count(/) 1实操心得category_path的正则r^[^/](/[^/]){0,4}$是经过血泪教训的。最初只写str结果上游传入electronics//smartphone双斜杠导致下游split(/)得到空字符串引发IndexError。现在这个正则强制路径规范且category_level属性让业务代码无需重复解析。4.3 步骤3搭建特征工厂基础20分钟在features/__init__.py中实现特征注册器import importlib from pathlib import Path FEATURE_REGISTRY {} def load_features(): features_dir Path(__file__).parent for py_file in features_dir.glob(*.py): if py_file.name in [__init__.py, base.py]: continue module_name ffeatures.{py_file.stem} module importlib.import_module(module_name) for attr_name in dir(module): attr getattr(module, attr_name) if hasattr(attr, _feature_metadata): FEATURE_REGISTRY[attr._feature_metadata[name]] attr # 装饰器 def feature(version: str, depends_on: list None): def decorator(func): func._feature_metadata { name: f{func.__name__}_v{version.replace(., _)}, version: version, depends_on: depends_on or [] } return func return decorator避坑提示importlib.import_module()必须用绝对导入否则在Docker容器内会报ModuleNotFoundError。我们吃过亏——本地开发用相对导入from .user_activity import ...没问题但容器里Python path不同必须统一用importlib.import_module(features.user_activity)。4.4 步骤4编写第一个可测试特征25分钟创建features/user_order_count.pyfrom features.base import execute_sql # 封装好的DB连接 from features import feature feature(version1.0.0, depends_on[user_orders]) def user_total_order_count(user_id: str) - int: 用户历史总订单数含取消订单 sql fSELECT COUNT(*) FROM orders WHERE user_id {user_id} result execute_sql(sql) return int(result[0][0]) if result else 0 # 测试文件 features/test_user_order_count.py def test_user_total_order_count(): # Mock DB返回固定值 with patch(features.base.execute_sql) as mock_sql: mock_sql.return_value [(123,)] assert user_total_order_count(U123) 123 # 测试空用户 with patch(features.base.execute_sql) as mock_sql: mock_sql.return_value [] assert user_total_order_count(U999) 0关键细节测试必须覆盖user_id不存在的场景返回0而非None因为下游模型可能用np.log1p()None会报错。我们曾因此在线上引发ValueError: cannot convert float NaN to integer。4.5 步骤5构建最小可行模型30分钟在models/resnet_classifier.py中定义模型import torch import torch.nn as nn from torchvision.models import resnet18 class ProductClassifier(nn.Module): def __init__(self, num_classes: int 1000): super().__init__() self.backbone resnet18(pretrainedFalse) # 关键不加载预训练权重 self.backbone.fc nn.Linear(512, num_classes) def forward(self, x): return self.backbone(x) def save(self, path: str): torch.save({ state_dict: self.state_dict(), num_classes: self.backbone.fc.out_features, timestamp: torch.datetime.now().isoformat() }, path)为什么禁用pretrainedTrue因为预训练权重是针对ImageNet的而我们的商品图数据分布完全不同。强行加载会导致前几层梯度爆炸。我们实测从零初始化训练在自有数据上收敛更快且最终准确率高1.2%。save()方法里存timestamp是为了在CI中做版本校验——如果两个commit的模型文件timestamp相同说明没重新训练。4.6 步骤6实现模型训练流水线45分钟Makefile中定义训练目标# Makefile TRAIN_DATA_PATH ? ./data/train MODEL_OUTPUT ? ./models/checkpoints/latest.pth train: python -m torch.distributed.launch \ --nproc_per_node2 \ --master_port29500 \ trainer.py \ --data_dir $(TRAIN_DATA_PATH) \ --output_dir $(MODEL_OUTPUT) \ --epochs 50 \ --batch_size 64 \ --lr 0.01 # trainer.py关键逻辑 def main(): # 1. 初始化分布式训练 dist.init_process_group(backendnccl) # 2. 加载数据用我们自己的Dataset非torchvision dataset ProductDataset(data_dirargs.data_dir, contractProductItem) # 3. 模型并行化 model ProductClassifier(num_classesdataset.num_classes) model DDP(model.cuda(), device_ids[args.local_rank]) # 4. 训练循环省略细节 for epoch in range(args.epochs): train_one_epoch(model, dataloader, optimizer) if args.rank 0: # 主进程保存 save_checkpoint(model, epoch)经验之谈ProductDataset必须继承torch.utils.data.Dataset但重写__getitem__时加入契约校验def __getitem__(self, idx): raw_data self._load_raw(idx) # 从磁盘读取 try: validated ProductItem(**raw_data) # 强制校验 return self._preprocess(validated) # 返回tensor except ValidationError as e: # 记录错误样本ID供数据团队修复 log_error(fInvalid sample {idx}: {e}) raise SkipSampleError() # 跳过此样本不停止训练4.7 步骤7导出ONNX并验证20分钟export_onnx.pyimport torch import onnx from models.resnet_classifier import ProductClassifier model ProductClassifier(num_classes1000) model.load_state_dict(torch.load(./models/checkpoints/latest.pth)[state_dict]) model.eval() # 创建dummy input必须匹配训练时的预处理 dummy_input torch.randn(1, 3, 224, 224) # 导出ONNX关键参数 torch.onnx.export( model, dummy_input, ./models/exported/model.onnx, opset_version12, # 不要用最新opset兼容性差 input_names[input], output_names[output], dynamic_axes{ input: {0: batch_size}, output: {0: batch_size} } ) # 验证ONNX onnx_model onnx.load(./models/exported/model.onnx) onnx.checker.check_model(onnx_model) # 必须通过血泪教训opset_version12是经过验证的黄金版本。我们试过opset15结果ONNX Runtime在某些GPU上报Unsupported operator: Resize。dynamic_axes必须声明否则Triton无法处理变长batch。4.8 步骤8编写Triton自定义Backend35分钟services/triton_backend/1/model.pyimport numpy as np import torch from models.resnet_classifier import ProductClassifier class TritonPythonModel: def initialize(self, args): self.model ProductClassifier(num_classes1000) self.model.load_state_dict( torch.load(/models/resnet_classifier/1/model.pth)[state_dict] ) self.model.eval() self.model.cuda() # 必须cuda() def execute(self, requests): responses [] for request in requests: # Triton传入的是numpy array input_array torch.from_numpy(request.input(input)).cuda() with torch.no_grad(): output self.model(input_array) # 转回numpyTriton要求 output_np output.cpu().numpy() responses.append( pb_utils.InferenceResponse( output_tensors[ pb_utils.Tensor(output, output_np) ] ) ) return responses核心配置config.pbtxtname: resnet_classifier platform: pytorch_libtorch max_batch_size: 32 input [ { name: input data_type: TYPE_FP32 dims: [3, 224, 224] } ] output [ { name: output data_type: TYPE_FP32 dims: [1000] } ]注意platform: pytorch_libtorch表示用LibTorch C API比pytorch_python快3倍。max_batch_size: 32不是拍脑袋而是根据GPU显存计算A100 40GB模型参数约45MBbatch_size32时显存占用≈2.1GB留足余量。4.9 步骤9部署K8s服务25分钟infra/k8s/deployment.yamlapiVersion: apps/v1 kind: Deployment metadata: name: resnet-classifier spec: replicas: 2 selector: matchLabels: app: resnet-classifier template: metadata: labels: app: resnet-classifier spec: containers: - name: triton-server image: nvcr.io/nvidia/tritonserver:23.08-py3 ports: - containerPort: 8000 name: http - containerPort: 8001 name: grpc resources: limits: nvidia.com/gpu: 1 memory: 8Gi requests: nvidia.com/gpu: 1 memory: 6Gi livenessProbe: httpGet: path: /v2/health/live port: 8000 initialDelaySeconds: 60 periodSeconds: 30 readinessProbe: httpGet: path: /v2/health/ready port: 8000 initialDelaySeconds: 120 # 关键给模型加载留足时间 periodSeconds: 10为什么initialDelaySeconds120因为A100加载1.2GB模型权重需要约90秒。我们用kubectl logs -f观察过TRITONSERVER_MODEL_REPOSITORY挂载后Triton日志显示Loading model resnet_classifier到Successfully loaded平均耗时112秒。设成60秒会频繁重启。4.10 步骤10实现三层健康检查15分钟在services/health_check.py中from fastapi import FastAPI, HTTPException from services.triton_client import TritonClient app FastAPI() app.get(/health/infrastructure) def infra_health(): return {status: ok, timestamp: datetime.now().isoformat()} app.get(/health/model) def model_health(): try: client TritonClient(localhost:8000) # 发送最小样本 dummy_input np.random.rand(1, 3, 224, 224).astype(np.float32) result client.infer(resnet_classifier, dummy_input) if result.shape ! (1, 1000): raise Exception(Wrong output shape) return {status: ok, latency_ms: 12.5} except Exception as e: raise HTTPException(status_code503, detailfModel error: {e}) app.get(/health/business) def business_health(): try: # 调用真实业务逻辑 response requests.post(http://localhost:8000/v2/models/resnet_classifier/infer, json{ inputs: [{name: input, shape: [1,3,224,224], datatype: FP32, data: [0.5]*3*224*224}] }) if response.status_code ! 200: raise Exception(Business logic failed) return {status: ok, business_metric: valid_inference} except Exception as e: raise HTTPException(status_code503, detailfBusiness error: {e})部署时绑定在K8s Service中readinessProbe指向/health/business而非/health/infrastructure。这样即使Triton进程活着但业务逻辑异常也会自动摘除流量。4.11 步骤11配置Prometheus监控20分钟infra/prometheus/config.ymlscrape_configs: - job_name: triton-server static_configs: - targets: [triton-service:8002] # Triton暴露metrics端口8002 metrics_path: /metrics - job_name: gateway static_configs: - targets: [gateway-service:8000] metrics_path: /metrics # 自定义指标在gateway代码中 from prometheus_client import Counter, Histogram PREDICTION_COUNTER Counter( prediction_total, Total number of predictions, [model_name, status] # status: success/fail ) PREDICTION_LATENCY Histogram( prediction_latency_seconds, Prediction latency in seconds, [model_name], buckets[0.01, 0.05, 0.1, 0.2, 0.5, 1.0, 2.0] )关键实践PREDICTION_COUNTER的status标签必须包含fail且fail原因要细分fail: timeout,fail: oom,fail: invalid_input。这样在Grafana中能快速定位瓶颈——如果是timeout突增说明模型推理慢如果是oom突增说明batch_size配置错误。4.12 步骤12运行端到端测试10分钟tests/e2e_test.pyimport pytest import requests def test_full_pipeline(): # 1. 准备测试数据符合ProductItem契约 test_data { sku_id: TEST123, title: test product, price_cny: 99.99, category_path: electronics/smartphone, update_time: 2024-01-01T00:00:00Z } # 2. 调用特征工厂 from features.user_order_count import user_total_order_count feature_val user_total_order_count(U001) # 3. 调用模型服务 response requests.post( http://localhost:8000/v2/models/resnet_classifier/infer, json{inputs: [{name: input, shape: [1,3,224,224], datatype: FP32, data: [0.0]*3*224*224}]} ) # 4. 断言 assert response.status_code 200 assert output in response.json() assert len(response.json()[output]) 1000 print(✅ End-to-end pipeline passed!) if __name__ __main__: test_full_pipeline()CI集成在GitHub Actions中这个测试是deployjob的前置条件。只有pytest tests/e2e_test.py通过才会执行kubectl apply -f infra/k8s/。我们曾因此拦截过一次灾难某次模型导出脚本bug导致ONNX输出维度是[1, 1001]而非[1, 1000]e2e测试直接失败避免了上线后大面积报错。5. 那些没人告诉你的实战陷阱与破局技巧5.1 陷阱一
返回列表