云端训推优化负责人 北京、深圳 全职 智能制造 / 工业互联网 / 工业自动化 职位描述 负责云端训练与推理优化团队整体建设和管理,对模型训推效率负责;负责大模型、VLM / VLA、机器人模型等训练任务的性能优化,包括:训练吞吐、GPU 利用率、Tensor Core 利用率、显存利用率、多机多卡扩展效率负责分布式训练优化,包括 Data Parallel、Tensor Parallel、Pipeline Parallel、FSDP / ZeRO 等方案设计与落地;负责模型推理优化,包括算子优化、KV Cache、Batching、量化、并行推理及服务化性能优化;针对模型结构和训练流程进行 Profiling,定位通信、计算、显存、IO 等性能瓶颈,并推动系统性优化;负责训练/推理框架和基础组件建设,沉淀可复用的性能优化能力;与模型算法团队协作,在保证模型效果的前提下优化模型结构、训练策略和推理方案;与智算平台团队协作,提升 GPU 集群整体资源利用率,降低单位训练与推理成本;建立训推性能指标体系和 Benchmark,持续跟踪不同模型、硬件和框架下的性能表现。 职位要求 5 年以上深度学习系统、训练加速、推理优化、AI Infra 或相关工作经验,有团队管理经验;熟悉 PyTorch 及主流深度学习训练框架,理解计算图、Autograd、CUDA 执行模型;熟悉大规模分布式训练,有 NCCL、FSDP、DeepSpeed、Megatron-LM
NVIDIA is leading company of AI computing. At NVIDIA, our employees are passionate about AI, HPC , VISUAL, GAMING. Our SA team is more focusing to bring NVIDIA new technology into difference industries. We help to
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential
We are now looking for an AI Developer Technology Engineer Interns. Intelligent machines powered by AI computers that can learn, reason and interact with people are no longer science fiction. Today, a self-driving car can meander
NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 30 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential
Intelligent machines powered by AI computers that can learn, reason and interact with people are no longer science fiction. Today, a self-driving car can meander through a country road at night and find its way. An
NVIDIA are seeking Solution Architects with specialized expertise in training and deploying Large Language Models (LLMs), implementing RAG workflows, and agentic inference. You will leverage the full NVIDIA software & hardware ecosystem to design, optimize, and
NVIDIA has been defining computer graphics, PC gaming, and accelerated computing for more than 25 years. With an outstanding legacy of innovation, driven by phenomenal technology, and extraordinary people, NVIDIA is looking for a strong technical
NVIDIA is the world leader in computer graphics, artificial intelligence, and accelerated computing. For over 30 years, we have been at the forefront of research and engineering around the greatest advances in technology. Our history of
At NVIDIA, we’re solving the world’s most ambitious problems with our groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. We are looking for a Developer Relations Manager to work with China industrial/research community to integrate
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential
NVIDIA is currently seeking a Solutions Architect for High-Performance Databases! Would you enjoy researching new algorithms and memory management techniques to accelerate databases on modern computer architectures? Do you like investigating hardware and system bottlenecks, and
NVIDIA is leading company of AI computing. At NVIDIA, our employees are passionate about AI, HPC , VISUAL, GAMING. Our SA team is more focusing to bring NVIDIA new technology into difference industries. We help to
We are now looking for an AI Developer Technology Engineer Intern. Intelligent machines powered by AI computers that can learn, reason and interact with people are no longer science fiction. Today, a self-driving car can meander
Join NVIDIA’s Cosmos Lab Infrastructure team to develop training and post-training systems for advanced Physical AI models, including world foundation models and robot policies. Our infrastructure connects training, inference, and evaluation with simulation and real-world robot
(高级)异构计算工程师 北京、杭州 全职 智能制造 / 工业互联网 / 工业自动化 职位描述 1、与算法、FPGA等团队紧密配合,完成深度学习系统在异构平台的高性能实现;2、参与异构平台深度学习推理引擎的研发,能够适配CPU、GPU、FPGA、NPU等平台,向上提供统一调用接口,内部完成高效、稳定的实现。 职位要求 1、本科及以上学历,计算机相关专业,三年以上相关工作经验;2、熟练掌握计算机体系结构相关知识,对各异构平台的优缺点有深入见解;3、熟悉C++编程,有良好的编程习惯和架构思想;4、良好的学习能力和强烈的责任心,良好的沟通和团队合作精神。加分项:1、 熟悉CUDA,OpenCL等异构编程者优先;2、熟悉编译原理、LLVM者优先;3、熟悉现代C++标准和泛型编程者优先;4、熟悉主流深度学习框架如Pytorch、MxNet、TensorFlow、TensorRT的一种或多种。 投递...
At NVIDIA, we’re solving the world’s most ambitious problems with our groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. We are looking for a Developer Relations Manager to work with China industrial/research community to integrate
At present, NVIDIA is exploring the immense possibilities of AI to define the forthcoming era of computing. Our division develops compiler technologies within the CUDA software stack that support the transformation of high-level parallel programs into efficient
端到端自动驾驶算法部署工程师 杭州、北京、武汉 全职 研发 - 算法 职位描述 1. 支持端到端模型导出适配和部署加速;2. 负责支持模型前后处理的板端算法实现和加速;3. 负责算法模块全流程的性能profiling,针对热点函数实现板端异构加速;4. 负责模型分布式训练加速,满足自动驾驶模型快速迭代。 职位要求 1. 硕士及以上学历,计算机相关专业,三年以上相关工作经验;2. 熟练掌握C++编程,熟悉常用linux工具,有良好的编程习惯;3. 深入理解pytorch等神经网络模型训练过程,熟悉常用分布式训练加速框架,掌握训练加速调优手段;4. 了解主要算子的量化方法,熟悉PTQ/QAT 进行加速推理的流程;5. 熟悉cuda并行加速和cuda算子开发;6. 英伟达、地平线等AI芯片平台有模型部署经验者优先;7. 良好的学习能力和强烈的责任心,良好的沟通和团队合作精神。 投递...
算子开发工程师 北京、深圳 全职 本科及以上 1-3 年 职位描述 1. 面向国内外多种 GPU 或其他 AI 算力进行高性能算子开发与优化。2. 面向实际模型与实际硬件平台,进行算子与推理引擎的综合优化。 职位要求 1. 计算机科学或相关领域硕士研究生及以上学历。2. 具有良好的学习能力和团队合作精神。3. 掌握 Python、C++、C 等通用编程语言,以及 CUDA 或 AscendC 等算子编程语言。4. 具备 GPU 等加速卡的算子优化或算子自动编译优化经验。5. 具备量化矩阵乘法、Flash Attention、计算通信融合算子等复杂算子优化经验者优先。6. 熟悉 PyTorch 框架、熟悉大模型完整推理流程者、熟悉算子与推理引擎联合优化者优先。7. 熟悉国内外最新 AI 算力硬件体系结构及其优化方法者优先。 投递...