Summary
At Apple, the AIML – On-Device Machine Learning group is responsible for accelerating the creation of amazing on-device ML experiences, and we are looking for a tenured software engineer to help define and implement features that accelerate and compress large state of the art (SoTA) models (e.g., LLMs) in our on-device inference stack. We are a dedicated team working on ground breaking technology in the field of natural language processing, computer vision and artificial intelligence. We are designing, developing, and optimizing large-scale language/vision/multi-modal models that power on-device inference capabilities across various Apple products and services. This is a unique opportunity to work on powerful new technologies and contribute to Apple’s ecosystem, with a commitment to privacy and user experience impacting millions of users worldwide.
Are you someone who can write high-quality, well-tested code and collaborate cross-functionally with partner HW, SW and ML teams across the company? If so, come join us and be part of the team that is helping Machine Learning developers innovate and ship enriching experiences on Apple devices!
Key Qualifications
5+ years proven programming skills using standard ML tools such as C/C++, CUDA/Metal, PyTorch, Tensorflow
Hands-on experience working on LLVM, compiler technologies, optimization techniques like quantization and sparsity-induction is a huge plus
Solid understanding of state-of-the-art DNN optimization techniques and how they translate to hardware acceleration architectures, and a general ability to reason about system performance (compute/memory) tradeoffs
Experience building APIs and/or core components of ML frameworks and strong attention to detail
Capacity to iterate on ideas, work with a variety of partners from all parts of the stack – from Apps to Compilation, HW Arch, and Power/Performance analysis
Excellent problem-solving (e.g. via building forward-looking prototype systems), critical thinking, strong communication, and collaboration skills
Description
As a member of this team, the successful candidate will:
– Build features for our on-device inference stack to support the most relevant accuracy preserving, general purpose techniques that empower model developers to compress and accelerate SoTA models (e.g., LLMs) in apps
– Convert models from a high-level ML framework to a target device (CPU, GPU, Neural Engine) for optimal functional accuracy and performance. Diagnose performance bottlenecks and work with HW Arch teams to co-design solutions that further improve latency, power, and memory footprint of neural network workloads
– Analyze impact of model optimization (compression/quantization etc) on model quality by partnering with modeling and adaptation teams across diverse product use cases. In this role, you will focus on optimizing our software stack for efficient execution on Apple GPUs, ANEs and CPUs
Education & Experience
Bachelor’s, Master’s, or PhD in Computer Science, Machine Learning or a related field
Additional Requirements
Project Role : AI / ML EngineerProject Role Description : Develops applications and systems that utilize AI tools, Cloud AI...
How to applyThe Position A healthier future. It’s what drives us to innovate. To continuously advance science and ensure everyone has access...
How to applySummary Imagine what you could do here. At Apple, great ideas have a way of becoming great products, services, and...
How to apply工作内容: 在Zoom 2.0时代,我们正在从-款优秀的现象级视频会议产品,逐步升级成协同办公平台,产品横跨了视频会议,语音电话,呼叫中心,智能助手,办公文档等.我们是Zoom AI基础设施团队,致力于为Zoom提供统-可靠新进的AI基础设施平台,Zoom AI正处于快速发展阶段,欢迎加入Zoom AI基础架构团队,您将和我们-道: 1,负责参与AI服务框架的研发,AI模型推理的评估,压测,调优. 2,负责参与建设AI流量网关,实现可编排可插拔的AI业务流水线,封装各种AI技能所需的SDK等. 3,负责参与AI基础调度能力的建设,包括GPU实例生命周期管理,GPU实例编排调度等,为业务提供-个满足企业级稳定性和性能要求的AI专用调度平台层. 岗位要求: 1,计算机相关专业本科以及以上学历,具备丰富的软件开发经验;积极向上,主动性强,有想法,抗压能力强,良好的沟通能力. 2,出色的业务架构能力,能应对复杂庞大的业务流量,提供高稳定性,高可用的保障. 3,熟悉AI模型研发流程,对流程中涉及到的技术框架有较为深入的了解. 4,了解分布式系统,调度,容器相关领域技术,熟悉Kubernetes/docker/Yarn等原理与实现,能基于通用方案进行二次研发. 5,精通任-主流公有云平台(AWS,Azure,GCP,Ali等)结构及技术特性,主流虚拟化技术,对IaaS / PaaS平台构架有深层次理解及技术洞察,有大型云平台系统架构设计经验. 加分项: 1,有高性能,高并发,高容错的服务开发经验者优先. 2,热于探索新技术,对于云原生/AI/大模型原理有深刻理解,熟悉AI相关编程语言(Python等)者优先....
How to applyJob Overview: The Cambridge HPC AI Technologist is a field-based consultant that builds end-to-end research computing system solutions. You will...
How to applyDo you want to take the first step in making Filipinos’ lives better everyday? Here in GCash we want to...
How to apply