【27届校招】SRE(系统可靠性工程师) 上海 正式 技术类 2027届秋季校园招聘项目 职位描述 1、负责公司业务系统运维工作,提升业务稳定性和工程效率,与业务方保持高效沟通,建立良好合作关系;2、负责应用上线评审、上线交付、配置变更、状态监控、容量管理、故障应急响应工作;3、参与业务服务端架构的高可用设计和性能优化,保证高效、可靠的业务迭代;4、负责线上重大问题排查,紧急事故处理,后续事故分析与优化;5、负责应用故障演练、应急预案、SOP手册编写工作,确保故障时业务能快速恢复;6、负责应用高可用建议及管理,包括限流、降级,容错、容灾,同城多活,确保应用质量;7、建立SLA评估标准,计算故障对SLA影响,并对SLA后续改进措施进行跟进;8、负责运维规范、流程文档编制,并将其工具化、平台化,确保运维安全,提升运维效率;9、探索 AI 在稳定性领域的应用,包括辅助故障根因分析、告警降噪运维工作智能化。 职位要求 1、2027届毕业生,本科以上学历,计算机类、软件类、通信类等相关专业优先;2、扎实的编程基础,至少掌握一门开发语言,熟悉Golang, python优先;3、对主流操作系统熟悉比如Linux,Debian等;4、熟悉容器技术、云原生;5、具备使用 AI 编程工具(如 Claude Code、Cursor)辅助开发和调试的实践经验;6、了解大模型/Agent 基本原理,对 AIOps、混沌工程自动化等方向有一定了解或探索优先。 投递...
SRE工程师 Shanghai Experienced Outsourced Responsibilities Vast 是一家专注于 3D AI 的独角兽公司,致力于构建能够理解和生成 3D 世界的基础模型。公司获得 NVIDIA、Lux Capital 等顶级机构投资,产品覆盖 3D 内容生成、机器人仿真训练、空间计算等前沿领域。SRE 团队是支撑这一切稳定运转的基础——你将直接参与构建服务于大规模 3D AI 模型训练与推理的基础设施体系。职位要求1、设计并维护 AWS / 阿里云的多账号体系(Organization / Landing Zone),规划账号结构与边界,制定并落地 IAM 权限策略,推行最小权限原则,管理角色、策略与 SCP2、建立云资源全生命周期管理规范,包括 Tagging 标准、资源审计与清理流程,推动 FinOps 实践:成本归因、预算告警、RI / SP 购买策略、Spot 实例优化3、通过
SummaryDo you want to help build some of the largest and most consequential enterprise and customer technology systems in the world? Join Apple’s Information Systems and Technology (IS&T) organization. IS&T is the engine behind everything Apple
SummaryThe people here at Apple don’t just build products — we craft the kind of wonder that’s revolutionized entire industries. It’s the diversity of those people and their ideas that supports the innovation that runs through
SummaryWe are hiring an SRE to own reliability, performance, data freshness, and operational readiness for our production ETL platform. The platform supports data ingestion, transformation, and loading workflows across Kubernetes-based environments — including Airflow-based loader jobs and
SummaryThe AiDP Data Platforms team builds and operates data platforms at scale on the Cloud, helping Apple process, store, and access petabytes of data. We’re seeking an SRE to own the reliability, performance, and operability of our
Job Description Have clear and solid relationships with software development departments. Plan and document work and projects. Build and continuously optimize CI/CD process and streamline automation effort for server provisioning and applications deployment. Build a resilient
Job Description Have clear and solid relationships with software development departments. Plan and document work and projects. Build and continuously optimize CI/CD process and streamline automation effort for server provisioning and applications deployment. Build a resilient
资深云原生基础设施工程师(Kubernetes/云网络方向) 上海 全职 互联网 / 电子 / 网游 职位描述 岗位职责1. 负责公司云原生基础设施的架构治理、生产运维与稳定性保障,覆盖多云托管 Kubernetes 集群及相关网络、存储、调度、发布和故障恢复能力。2. 深入负责 Kubernetes 数据面与公有云基础设施,重点建设和治理 CNI、CSI、VPC、负载均衡、DNS、路由、安全组、ENI、NetworkPolicy等能力,解决跨节点、跨可用区及多云环境中的复杂网络与存储问题。3. 逐步承接业务运维侧的 VPC 网络职责,梳理从公网或内网入口、云 LB、VPC、节点到 Service、CNI及Pod的完整流量路径,建立自主排障和稳定性保障能力。4. 负责 ACK、TKE、EKS、GKE 等多云托管 Kubernetes 环境的运维、治理与优化,推动集群基线、容量管理、弹性扩缩容、发布回滚和资源治理标准化。5. 参与游戏业务的部署交付、线上保障和重大活动保障,承担必要的值班及应急响应,在复杂故障中负责核心定位与处置,推动问题复盘和闭环整改。6. 与平台工程、研发、安全等团队协作,完善监控、告警、日志、Event和链路追踪等排障闭环;推动多集群、安全治理及云原生能力的工程化落地。7. 沉淀云原生基础设施相关文档、SOP、排障手册和最佳实践,提升团队对云网络、容器网络及存储体系的接管能力。任职资格1. 全日制本科及以上学历,计算机相关专业,5 年及以上运维、SRE、平台工程或云原生基础设施经验,具备大型互联网、游戏或高并发业务生产经验。2. 具备扎实的 Linux 和 TCP/IP 基础,能够独立分析生产环境中的网络、存储、性能及稳定性问题。3. 具备完整的
SummaryImagine what you could do here. At Apple, new ideas have a way of becoming extraordinary products, services, and customer experiences very quickly. We are looking for great Systems Reliability Engineers to design, build and operate
Job Responsibilities: -Design and develop enterprise-grade AI application systems, driving the deep integration of AI into manufacturing and business operations. -Build systems—such as intelligent quality inspection, predictive maintenance, knowledge Q&A, and intelligent production scheduling—leveraging Large Language
Job Responsibilities: - Design and develop enterprise-level AI application systems to drive the deep integration of artificial intelligence into manufacturing and business scenarios. - Build systems such as intelligent quality inspection, predictive maintenance, knowledge Q&A, and
SummaryThe Operations Engineer in Crypto Services team manages key technical infrastructure. An ideal candidate will have experience in Systems Administration. The Operations Engineer will monitor infrastructure and application services and drive incident management. The Ops engineer
SummaryJoint Mobile Engineering Team (JMET) is a security engineering team that provides critical services for Apple across every product line. From manufacturing to customer-facing operations, the teams services span the entire lifecycle of most Apple hardware.
SummaryImagine what we could do together. At Apple, new ideas have a way of becoming excellent products, services, and customer experiences very quickly. Bring passion and dedication to your job and there’s no telling what you
运维开发工程师 武汉、西安、上海 校招 正式 运维类 2027届校园招聘计划 职位描述 岗位描述: 1、我们为小米的站点稳定性提供系统和工具,是稳定性工程专家;2、实践DevOps 和 Google SRE的工程思想;3、站在软件生命周期全局考虑如何保障系统健壮性,稳定性,并通过软件工程技术为站点保驾护航;4、站在用户/工程师的角度提供易用的工具和平台产品,不断完善使用体验;5、负责集团基础软件/工具/平台的设计和研发,维护小米各项业务的稳定运行;6、按照优秀的工程实践完成需求,设计,编码,测试,发布的软件开发流程;7、按照标准编写设计/开发/运维/用户文。 职位要求 任职要求: 1、熟练C/C++/Java/【Go】至少一种编程语言,有Ruby/Python/Bash经验更好;2、熟悉Linux/Unix系统/【K8s 生态】;3、熟悉软件开发方法/流程,理解软件设计,开发,测试,发布过程;4、持续学习,不断追求更好,不断挑战自己,包括: 产品/架构/方案/流程/编码;5、乐于分享,具备服务精神,良好的沟通能力和团队合作精神;6、优秀的分析和解决问题能力,勇于解决难题,有”问题到我为止“的精神。 投递...
SummaryImagine what we could do together. At Apple, new ideas have a way of becoming excellent products, services, and customer experiences very quickly. Bring passion and dedication to your job and there’s no telling what you
The role is based in the GMT+8 Timezone. We can engage you in Singapore, China or remotely if applicable. About Us DeGate is a self-custodial crypto wallet built to make earning on blockchains simple. Users deposit
Job Description Have clear and solid relationships with software development departments. Plan and document work and projects. Build and continuously optimize CI/CD process and streamline automation effort for server provisioning and applications deployment. Build a resilient