AWS EC2 Spot 实例抢单策略:基于 Auto Scaling Group 的成本优化(附失败重试脚本)
·
AWS EC2 Spot 实例抢单策略:基于 Auto Scaling Group 的成本优化
核心优化策略
-
混合实例组合
在 Auto Scaling Group (ASG) 中配置混合实例策略:mixed_instances_policy = { "SpotPercentage": 70, # 70% Spot实例 "OnDemandBaseCapacity": 2, # 至少2台按需实例保底 "InstanceTypes": ["c5.large", "m5.large", "r5.large"], # 多实例类型 "AllocationStrategy": "capacity-optimized" # 容量优化策略 }- 优势:平衡中断风险与成本节约(Spot 实例价格比按需低 60-90%)
- 关键参数:
capacity-optimized:自动选择中断率最低的实例池- 多实例类型:增加抢单成功率
- 按需保底:确保服务连续性
-
中断处理机制
Spot 实例中断时 ASG 自动响应:graph LR A[中断警告] --> B[2分钟缓冲期] B --> C[ASG 启动新实例] C --> D[负载均衡器流量转移]
失败重试脚本
import boto3
import time
def retry_spot_failures(asg_name, max_retries=3):
ec2 = boto3.client('ec2')
asg = boto3.client('autoscaling')
for attempt in range(max_retries):
# 检测失败的Spot请求
failed_requests = ec2.describe_spot_instance_requests(
Filters=[{'Name': 'state', 'Values': ['failed']}]
)
if not failed_requests['SpotInstanceRequests']:
print("无失败请求")
return True
print(f"检测到 {len(failed_requests)} 个失败请求,重试中...")
# 重新提交请求
for request in failed_requests:
ec2.request_spot_instances(
InstanceCount=1,
LaunchSpecification=request['LaunchSpecification'],
Type='persistent'
)
# 等待ASG稳定
time.sleep(120)
asg.enter_standby(AutoScalingGroupName=asg_name)
asg.exit_standby(AutoScalingGroupName=asg_name)
return False
# 使用示例
if retry_spot_failures("prod-asg"):
print("实例恢复成功")
else:
print("重试失败,触发告警")
关键优化指标
| 指标 | 优化前 | 优化后 |
|---|---|---|
| 计算成本 | $$C_{\text{ondemand}}$$ | $$0.3C_{\text{ondemand}}$$ |
| 中断率 | >15% | <5% |
| 恢复时间 | 手动操作 | <3分钟 |
实施步骤
- ASG 配置
aws autoscaling create-auto-scaling-group \ --mixed-instances-policy file://policy.json \ --capacity-rebalance - 中断预警处理
部署 CloudWatch 事件规则:{ "source": ["aws.ec2"], "detail-type": ["EC2 Spot Instance Interruption Warning"] } - 成本监控
使用 Cost Explorer 公式:
$$ \text{Savings} = \sum (\text{OnDemandPrice}_i - \text{SpotPrice}_i) \times t_i $$
最佳实践
- 选择多个可用区(AZ)分散风险
- 避免单一实例类型,建议至少 3 种备选类型
- 结合 Savings Plans 进一步降低按需实例成本
- 使用 EC2 中断预测工具(需启用 AWS Compute Optimizer)
通过此方案,典型业务可降低 60-70% 计算成本,中断率控制在 5% 以内。重试脚本应配合 CloudWatch 告警使用,当连续失败时触发人工介入。
更多推荐

所有评论(0)