Are you fascinated by the chance to make a real impact on millions of Ubuntu users? We are looking for a passionate Linux kernel engineer to join our team and help us bring Ubuntu to the next
2027校招-算子开发工程师(上海/北京/成都) 上海、北京、成都 校招 正式 职位描述 岗位职责- 负责公司自研 AI 加速芯片上的 算子库设计与实现,覆盖训练与推理核心算子(如 GEMM/Conv/Attention/LayerNorm/Softmax/Reduce/Scatter-Gather 等)。- 基于芯片计算单元、片上存储(SRAM/Cache)、HBM 带宽、NoC 拓扑、DMA/异步拷贝能力,进行 算子性能建模与瓶颈分析(roofline、访存/计算比、流水并行),制定优化策略。- 设计并落地 kernel 调度与 tiling 策略:多级分块、数据布局(layout/packing)、向量化/张量化、双缓冲/多缓冲、算子融合(fusion)等。- 建设算子库工程化体系:算子接口规范、版本管理、性能回归、单测/对齐验证、精度与数值稳定性、跨平台兼容- 参与端到端模型性能优化:对接 PyTorch/TensorFlow/ONNX Runtime 等前端(或公司自研框架),完成算子覆盖、性能调优和线上问题定位(profiling/trace)。 职位要求 (优秀应届可放宽)。- 计算机/电子/自动化/数学等相关专业;- 熟悉深度学习常见算子与数值实现(GEMM/Conv/Attention/Norm/Reduce),理解训练/推理的差异(激活、梯度、混合精度、量化)。- 扎实的 C/C++ 能力,良好的工程习惯;能熟练使用性能分析工具(profiling、trace、perf、火焰图或自研工具)。- 熟悉至少一种加速器编程模型或并行体系(CUDA/OpenCL/ROCm/Metal/TVM/Triton/oneDNN/CUTLASS 等)并能迁移到自研 ISA/Runtime。- 具备系统级性能意识:理解存储层次、带宽/延迟、cache/预取、NUMA、DMA、并行同步等,对算子优化有方法论。加分项-
嵌入式Linux底软工程师 北京 社招 全职 职位描述 1.负责机器人端侧计算平台的 Linux 底层软件开发与维护。2.负责系统启动、内核、文件系统、设备驱动及版本集成等工作。3.参与自研硬件的板级适配、设备点亮和软硬件联合调试。4.分析并解决系统实时性、数据链路性能和长期运行稳定性问题。5.配合硬件、上层软件及芯片厂商,推动底层问题定位和产品落地。 职位要求 1.计算机、电子、通信、自动化等相关专业,本科及以上学历。2.3年以上嵌入式 Linux、BSP或设备驱动开发经验。3.熟悉 C/C++,精通 Linux Kernel、Device Tree及驱动开发基础。4.有机器人、工业控制、智能设备或其他端侧芯片产品的实际开发经验和落地经验。5.具备较强的系统调试和问题定位能力,具备良好的工程意识、协作能力和问题闭环能力。加分项1.有端侧 Linux 产品从板级适配到稳定交付的完整经验。2.有系统实时性、Camera、多媒体或 MCU 协同开发经验。 投递...
Join our Team AI-Native Development Intern We build AI-native quality engineering and digital workforce solutions. You will work with experienced engineers to develop AI agents, automation tools, and intelligent workflows that improve software quality, delivery efficiency,
Canonicals Device Delivery Team works with tier-1 OEM and ODM customers to pre-load Ubuntu Desktop and Ubuntu Core, bringing Ubuntu directly to millions of users. As a Software Engineering Manager you will lead and manage the
云端训推优化负责人 北京、深圳 全职 智能制造 / 工业互联网 / 工业自动化 职位描述 负责云端训练与推理优化团队整体建设和管理,对模型训推效率负责;负责大模型、VLM / VLA、机器人模型等训练任务的性能优化,包括:训练吞吐、GPU 利用率、Tensor Core 利用率、显存利用率、多机多卡扩展效率负责分布式训练优化,包括 Data Parallel、Tensor Parallel、Pipeline Parallel、FSDP / ZeRO 等方案设计与落地;负责模型推理优化,包括算子优化、KV Cache、Batching、量化、并行推理及服务化性能优化;针对模型结构和训练流程进行 Profiling,定位通信、计算、显存、IO 等性能瓶颈,并推动系统性优化;负责训练/推理框架和基础组件建设,沉淀可复用的性能优化能力;与模型算法团队协作,在保证模型效果的前提下优化模型结构、训练策略和推理方案;与智算平台团队协作,提升 GPU 集群整体资源利用率,降低单位训练与推理成本;建立训推性能指标体系和 Benchmark,持续跟踪不同模型、硬件和框架下的性能表现。 职位要求 5 年以上深度学习系统、训练加速、推理优化、AI Infra 或相关工作经验,有团队管理经验;熟悉 PyTorch 及主流深度学习训练框架,理解计算图、Autograd、CUDA 执行模型;熟悉大规模分布式训练,有
Company: Qualcomm China Job Area:Engineering Group, Engineering Group Machine Learning Engineering General Summary: About us: We are Qualcomm AI Research that are advancing AI to make its core capabilities – perception, reasoning, and action – ubiquitous
This is an outstanding opportunity for switch SDK development engineer to join our multi-site team for switch and router related SW development. The successful candidate will collaborate closely with other development teams, arch and QA to
We are now looking for a Deep Learning Performance Software Engineer! We are expanding our research and development for deep learning. We seek excellent Software Engineers and Senior Software Engineers to join our team. We specialize
NVIDIA are seeking Solution Architects with specialized expertise in training and deploying Large Language Models (LLMs), implementing RAG workflows, and agentic inference. You will leverage the full NVIDIA software & hardware ecosystem to design, optimize, and
NVIDIA is building state-of-the-art accelerated computing platforms that know no boundaries. Our technology is essential for global innovators, scientists, researchers, and engineers empowering them to transform their boldest concepts into tangible outcomes. Our next-generation Infiniband, NVLink,
NVIDIA is currently seeking a Solutions Architect for High-Performance Databases! Would you enjoy researching new algorithms and memory management techniques to accelerate databases on modern computer architectures? Do you like investigating hardware and system bottlenecks, and
NVIDIA is leading company of AI computing. At NVIDIA, our employees are passionate about AI, HPC , VISUAL, GAMING. Our SA team is more focusing to bring NVIDIA new technology into difference industries. We help to
NVIDIA is seeking a passionate, world-class software engineer to join its Compute Developer Technology team(DevTech). Our team has over 150 engineers across Beijing, Shanghai, Shenzhen, Taipei, Seoul, and Sydney. We understand algorithms, GPU, and real-world applications.
Responsibilities Participate in the development and deployment of Generative AI use cases to improve engineering productivity, knowledge management, and business workflow automation. Design and implement AI-powered solutions using Microsoft Copilot, Copilot Studio, Glean, Vero, Power Automate,
About the Opportunity Personal Connect is recruiting a Firmware Engineer (Embedded Linux / SoC) on behalf of an innovative technology company operating at the intersection of embedded systems, hardware-software integration, and high-performance computing infrastructure. The company
About the Opportunity Personal Connect is recruiting a Senior Confidential Computing Infrastructure Engineer on behalf of an innovative technology company in Beijing. This is a highly specialized infrastructure engineering position focused on building secure environments for