AWS EC2 Spot 实例抢单策略:基于 Auto Scaling Group 的成本优化

核心优化策略
  1. 混合实例组合
    在 Auto Scaling Group (ASG) 中配置混合实例策略:

    mixed_instances_policy = {
        "SpotPercentage": 70,  # 70% Spot实例
        "OnDemandBaseCapacity": 2,  # 至少2台按需实例保底
        "InstanceTypes": ["c5.large", "m5.large", "r5.large"],  # 多实例类型
        "AllocationStrategy": "capacity-optimized"  # 容量优化策略
    }
    

    • 优势:平衡中断风险与成本节约(Spot 实例价格比按需低 60-90%)
    • 关键参数:
      • capacity-optimized:自动选择中断率最低的实例池
      • 多实例类型:增加抢单成功率
      • 按需保底:确保服务连续性
  2. 中断处理机制
    Spot 实例中断时 ASG 自动响应:

    graph LR
    A[中断警告] --> B[2分钟缓冲期]
    B --> C[ASG 启动新实例]
    C --> D[负载均衡器流量转移]
    

失败重试脚本
import boto3
import time

def retry_spot_failures(asg_name, max_retries=3):
    ec2 = boto3.client('ec2')
    asg = boto3.client('autoscaling')
    
    for attempt in range(max_retries):
        # 检测失败的Spot请求
        failed_requests = ec2.describe_spot_instance_requests(
            Filters=[{'Name': 'state', 'Values': ['failed']}]
        )
        
        if not failed_requests['SpotInstanceRequests']:
            print("无失败请求")
            return True
            
        print(f"检测到 {len(failed_requests)} 个失败请求,重试中...")
        
        # 重新提交请求
        for request in failed_requests:
            ec2.request_spot_instances(
                InstanceCount=1,
                LaunchSpecification=request['LaunchSpecification'],
                Type='persistent'
            )
        
        # 等待ASG稳定
        time.sleep(120)
        asg.enter_standby(AutoScalingGroupName=asg_name)
        asg.exit_standby(AutoScalingGroupName=asg_name)
        
    return False

# 使用示例
if retry_spot_failures("prod-asg"):
    print("实例恢复成功")
else:
    print("重试失败,触发告警")

关键优化指标
指标优化前优化后
计算成本$$C_{\text{ondemand}}$$$$0.3C_{\text{ondemand}}$$
中断率>15%<5%
恢复时间手动操作<3分钟
实施步骤
  1. ASG 配置
    aws autoscaling create-auto-scaling-group \
      --mixed-instances-policy file://policy.json \
      --capacity-rebalance
    

  2. 中断预警处理
    部署 CloudWatch 事件规则:
    {
      "source": ["aws.ec2"],
      "detail-type": ["EC2 Spot Instance Interruption Warning"]
    }
    

  3. 成本监控
    使用 Cost Explorer 公式:
    $$ \text{Savings} = \sum (\text{OnDemandPrice}_i - \text{SpotPrice}_i) \times t_i $$
最佳实践
  1. 选择多个可用区(AZ)分散风险
  2. 避免单一实例类型,建议至少 3 种备选类型
  3. 结合 Savings Plans 进一步降低按需实例成本
  4. 使用 EC2 中断预测工具(需启用 AWS Compute Optimizer)

通过此方案,典型业务可降低 60-70% 计算成本,中断率控制在 5% 以内。重试脚本应配合 CloudWatch 告警使用,当连续失败时触发人工介入。

更多推荐