NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential
NVIDIA is developing processor and system architectures that accelerate deep learning and high-performance computing applications. We are looking for an expert deep learning system performance architect to join our AI performance modelling, analysis and optimization efforts.
NVIDIA has continuously reinvented itself over two decades. NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning
端侧AI infra工程师 北京、上海 全职 职位描述 1.优化端侧微调+推理流程,加速端侧部署2.负责算法在真实业务场景中的落地,紧贴业务需求,不断改进算法在业务中的效果3.对接模型优化、模型部署、应用框架等多种能力,为算法交付提供有力支撑4.积极跟进AI学术界和业界的最新动态,优化内部算法模型,改进模型部署效率 职位要求 1.熟悉掌握常见有监督、无监督等算法模型的原理、优缺点、适用场景等基础知识;2.熟悉lerobot、gr00t等具身框架中的训练机制,包括但不限于BC,RL,HIL等3.熟悉Jax、Pytorch等主流深度学习框架,并有实际的模型训练、调优的项目经验;4.熟悉torch profiler,torch compile,tirtion,cudagraph等训练优化手段,并可以快速赋能训练过程5.熟悉transformer engine,model opt,torchao等算法优化框架,可以配合模型量化、剪枝、蒸馏等工作加速算法推理效果6.良好的沟通能力、解决问题能力;7.参与AI平台的研发,提升系统性能和稳定性;8.与团队合作,共同解决技术难题,推进项目进展。 投递...
AI编译器开发工程师 北京、上海 全职 职位描述 全栈开发自研编译器,辅助算法业务落地跟进算法交付,在规定时间内交付可部署的算法,并支持多芯片部署持续优化编译器各层能力,保持编译效果对齐业界SOTA 职位要求 5年以上AI编译经验,独立部署过vla,vlm类型模型;理解AI框架及常见的AI模型;熟练掌握python以及C/C++编程;对AI编译技术栈有深度理解,前端、优化、后端有基本认知熟悉transformers,hugginface,pytorch等框架的使用并了解算法搭建过程深度理解torch.compile, torch.trace, jax.jit过程原理,有能力进行AST语法树解析熟练使用TVM,TensorRT等框架并有实际部署经验了解算子优化与集成常用技术,可以使用算子注册加速模型 投递...
Be Part of an Amazing Story Macy’s is more than just a store. We’re a story. One that’s captured the hearts and minds of America for more than 160 years. A story about innovations and traditions…about