AI 平台高级 SRE 工程师 / DevOps 工程师 北京 全职 互联网 / 电子 / 网游 职位描述 1. 负责理想汽车 AI 平台大规模 GPU 集群、RDMA 网络及并行高速存储的稳定运行与日常运维,保障训练、推理等核心 AI 业务高效交付。 2. 负责大规模 GPU 集群在训练任务、资源调度、网络通信、存储访问等场景下的故障定位与问题解决,持续提升平台稳定性和可用性。 3. 深入理解 AI 训练与推理业务场景,参与 AI 平台在 Kubernetes 多集群、监控告警、日志检索、故障诊断等方向的云原生架构演进与落地。 4. 建设和完善
Job Description: • 具备大型数据平台及云环境(如Azure、AWS、Snowflake、Databricks)架构和管理经验,具备敏捷、DevOps和DataOps实践的实际知识。 • 熟练使用容器化和云自动化工具(Docker、Kubernetes、Terraform等),并对数据生命周期管理和数据治理最佳实践有深刻理解。 • 云计算平台(如AWS、Azure、GCP) • 基础设施即代码(IaC)工具(如 Terraform、CloudFormation) • Kubernetes和容器化(例如Docker) • CI/CD工具(例如,Jenkins、GitLab CI、GitHub Actions) • 脚本和编程语言(例如 Python、Go 语言) • 系统架构与设计 • 解决问题和故障排除 • 沟通与协作 At DXC Technology, we believe strong connections and community are
SRE实习生-2027届 南京 校招 实习 软件研发类 职位描述 1、负责小米各业务的SRE工作,比如AIoT、小米商城、互联网业务、小米手机核心应用、视频技术等2、工作涵盖容量管理、灾备管理、活动重保、日常Oncall、troubleshooting、业务巡检、故障预案、架构优化、技术运营等;3、与DEVS共同设计产品后端架构,实现分布式、全球集群化运维管理,制定并实施相关运维技术方案,确保服务高效、稳定的运行; 4、研发设计自动化运维工具,减少日常重复性工作,用DevOps工具化思维解决业务问题,提升运维效率; 5、通过技术手段进行成本控制及优化,通过工具化及流程提升服务管理效率。 职位要求 1、熟练至少一种编程语言:Go/Python/Bash/C/C++/Java;2、熟悉Linux/Unix系统;3、乐于分享、开源,具备服务精神,良好的沟通能力和团队合作精神;4、优秀的分析和解决问题能力,勇于解决难题,有”问题到我为止“的精神。 投递...
SRE工程师 热招 深圳 全职 研发 热招职位 职位描述 1、负责业务基础环境建设与维护,保障系统稳定运行;2、推进DevSecOps,协同安全部门落实安全要求与政策,提升企业安全水平;3、维护容器平台稳定性,完善监控与应急机制,确保故障快速定位与恢复;4、建设与优化流量治理、可观测和应急响应体系;5、通过自动化工具链与平台的构建,持续优化并提升交付效率与质量。 职位要求 1. 本科及以上学历,计算机科学、软件工程或相关专业优先;2. 3年以上系统运维/SRE/DevOps经验;有大规模容器集群管理或云环境项目经验者优先,制造业/全球分布式系统背景加分;3. 容器与编排:精通K8s,具备大规模集群管理、升级/迁移经验(如100+节点环境),熟悉Helm/Kustomize包管理;4、掌握安全集成方法,如容器安全基线(CIS Benchmarks)、CI/CD安全插件(SonarQube、Trivy)、配置管理安全(Secrets管理、RBAC);5、熟悉AWS(EKS/ECS)、阿里云(ACK)或GCP(GKE)等主流服务,包括云网络、存储与安全配置;有混合云部署经验优先;6、掌握至少一门脚本语言(如Python/Go/Shell),用于工具开发、API集成与自动化运维(如Ansible playbook、Python监控脚本);7、软性素质: - 高度责任心:对系统可靠性有强主人翁感,能主动识别风险并推动改进; - 沟通协作能力:跨团队(开发、安全、业务)高效协调,善于根因分析与问题复盘; - 应对挑战:高压环境下保持冷静,快速学习新技术,支持24/7 on-call(轮值响应)。 投递...
Job Summary As a Site Reliability Engineer (SRE) within Markets Technology, you will be responsible for improving the reliability, stability, performance, scalability, and operational resilience of the bank’s Markets applications and platforms. Working closely with development, infrastructure,
Job Summary As a Site Reliability Engineer (SRE) within Markets Technology, you will be responsible for improving the reliability, stability, performance, scalability, and operational resilience of the bank’s Markets applications and platforms. Working closely with development, infrastructure,
Job Summary As the Lead, Markets Site Reliability Engineering, you will define and drive the reliability engineering strategy across the Markets technology estate. You will lead a global team of Site Reliability Engineers, establishing SRE standards, governance,
Job Description Have clear and solid relationships with software development departments. Plan and document work and projects. Build and continuously optimize CI/CD process and streamline automation effort for server provisioning and applications deployment. Build a resilient
Some careers have more impact than others. If you’re looking for a career where you can make a real impression, join HSBC and discover how valued you’ll be. We are currently seeking an experienced professional to
Some careers have more impact than others. If you’re looking for a career where you can make a real impression, join HSBC and discover how valued you’ll be. We are currently seeking an experienced professional to
SummaryThe Insight team runs one of Apples most critical Big Data ecosystems — an exabyte-scale, highly-available infrastructure that underpins manufacturing operations for every Apple product, globally. Every iPhone, iPad, and Mac has touched our systems. We
资深运维开发工程师 热招 上海 全职 研发 职位描述 我们正在寻找一位兼具 稳定性治理能力 与 运维开发能力 的资深工程师,加入云端 SRE 团队,负责支撑业务增长阶段下的多云、多集群云原生基础设施稳定运行与持续优化。你将面向业务增长带来的稳定性、性能、容量和成本挑战,参与 Kubernetes 集群治理、Elasticsearch 等关键基础组件优化、线上故障治理、容量规划和变更风险控制。同时,你也将推动自动化运维平台和工具链建设,将线上问题沉淀为平台能力、工程规范和长期机制,提升研发、数据、安全、合规等团队的协作效率。1. 稳定性治理:负责云端基础设施及关键基础组件的稳定性建设,定位并解决线上性能瓶颈、容量风险和可用性问题,保障业务系统稳定运行;2. 性能优化:针对 Elasticsearch 等核心组件开展性能调优、容量评估、资源治理和架构优化,提升系统吞吐、查询效率和服务可靠性;3. 云原生基础设施:负责 Kubernetes 集群及 CNCF 云原生生态组件的日常运维、架构优化和稳定性提升,支撑并保障多个 Kubernetes 集群的可靠运行;4. 多云平台治理:参与阿里云 ACK、AWS EKS、GCP GKE 等多云托管 Kubernetes 环境的运维、治理和优化,提升多云环境下的可观测性、弹性、成本效率和运维一致性;5. 故障与变更管理:负责线上告警处理、故障应急、根因分析、复盘改进和生产变更管理,建立可持续的稳定性改进机制;6. 自动化与平台建设:开发和维护自动化运维平台、工具链和流程系统,提升发布、变更、巡检、告警、权限、资源交付等环节的自动化水平;7. 跨团队协作:与后端研发、数据、安全、合规等团队紧密协作,推动基础设施问题定位、流程规范、权限治理、合规要求和稳定性改进落地。
Minimum qualifications: Bachelor’s degree in Computer Science, a related technical field, or equivalent practical experience. 15 years of experience in a customer-facing technical role. Experience with core cloud computing concepts and enterprise architecture principles, (e.g., networking,
SummaryPeople at Apple don’t just build products — they craft experiences our customers love and depend on. Apple Services Engineering (ASE) builds and supports the systems that make many of these daily experiences possible. If you’ve
Who we are looking for We are looking for a senior hands-on individual contributor to provide critical production and operational support for traders, portfolio managers, and investment teams. The role requires strong knowledge of the trading
Position Summary State Street is seeking an experienced and transformational technology leader to lead the Hangzhou-based IMSWest Remediation Engineering team. This role will be accountable for building and leading a high-performing engineering team responsible for delivering
State Street is seeking a talented and motivated Software Engineer to join the IMSWest Remediation Engineering team in Hangzhou. This role will be responsible for designing, developing, testing, and supporting remediation solutions that improve platform stability,
Position Summary State Street is seeking a highly motivated Assistant Vice President (AVP) to lead application development initiatives within the IMSWest Remediation Program. This role will be responsible for driving the design, development, and delivery of
SummaryWe are hiring a Site Reliability Engineer to Build, support and improve Keystone, an ETL application/platform operating in the China region. Keystone supports data ingestion, transformation, and loading workflows across Kubernetes-based environments, including Airflow-based loader jobs
State Street is seeking a highly motivated AWS Federated Cloud Platform Engineer. This role is responsible for the design, development, enhancement, and operational support of the enterprise AWS Federated Landing Zone platform especailly on AWS compute