Applied Science Intern - Clipchamp
New South Wales, Australia · Victoria, Australia · Queensland, Australia · Brisbane, Australia · Melbourne, VIC, Australia · Sydney, NSW, Australia · Baw Baw, VIC, Australia · Cook, QLD, Australia · Hay, NSW, Australia
We’re looking for an Applied Science PhD student ready to bring their skills and ideas into the real world and gain hands-on experience working at the intersection of computer vision, generative AI and agentic models.
Our 12-week full-time Summer Internship gives you the opportunity to work alongside experienced applied scientists and engineers on real-world projects, exploring how AI can transform the way people create, edit and understand video.
At Clipchamp, you’ll join a global, cross-disciplinary team building video experiences that empower anyone to tell stories worth sharing. As an Applied Science PhD Intern, you’ll collaborate with engineers, product managers, designers and researchers to help reimagine how video is created, edited and understood.
You’ll design and evaluate multimodal models that see and reason over video and build agentic systems that can plan, call tools and complete real editing tasks on a user’s behalf. You’ll have access to cutting-edge generative AI models, LLMs, AI tooling and research, while gaining hands-on experience translating frontier technology into real product experiences.
Throughout the internship, you’ll have the opportunity to deepen your expertise in large-scale multimodal and agentic systems and strengthen your skills in experimental design, evaluation and collaborative research engineering, while making a meaningful contribution from day one.
About the Program: This program is open to university students who will graduate between August 2027 and July 2028 and can commit to a 12-week full-time internship between late November 2026 and February 2027.
This is an in-person internship based in one of our main Australian offices: Brisbane, Sydney or Melbourne. You’ll work alongside the team in the office and be part of the day-to-day Clipchamp experience. Please note: remote internship arrangements are not available for this role.
As part of Microsoft’s global intern community, you’ll also have opportunities to build your network, explore your interests and learn from others throughout your internship.
Responsibilities
- Research, prototype and evaluate computer vision and multimodal models for video understanding.
- Design and build agentic workflows in which language and multimodal models plan, invoke tools, and act over an editing timeline, then measure their reliability, latency and quality against real product scenarios.
- Translate ambiguous product goals into well-defined machine learning tasks, and design experiments, baselines and metrics that allow rapid iteration and optimization.
- Fine-tune, adapt and benchmark state-of-the-art foundation models (for example, vision-language models, diffusion models and LLM-based agents) on domain data, under the guidance of a senior scientist.
- Prepare and curate datasets for training and evaluation, reviewing data for quality and technical constraints, and documenting the actions taken to address data quality issues.
- Implement prototypes of scalable AI components and contribute to code reviews, analysis and technical documentation.
- Build an understanding of the broader research area and industry trends, and share your findings with the team through demos, write-ups and presentations.
Qualifications
Required Qualifications
- Currently pursuing a Doctorate degree in Computer Science, Applied Science, Statistics, or a related field, with a research focus in computer vision, machine learning or multimodal AI.
- Must have at least 1 semester/term remaining following the completion of the internship.
- Hands-on experience building and training deep learning models in Python with frameworks such as PyTorch, Hugging Face Transformers or Diffusers.
- Demonstrated experience in computer vision or video understanding, evidenced by research publications, open-source contributions or substantial project work.
Preferred Qualifications
- Experience with agentic AI systems, including tool use and function calling, planning and reasoning, multi-step orchestration, or agent evaluation frameworks.
- Familiarity with state-of-the-art architectures and techniques, such as transformers, attention mechanisms, vision-language models, diffusion models, transfer learning and parameter-efficient fine-tuning.
- Publication record at major AI conferences, including NeurIPS, ICLR, CVPR, ICML, ACL, EMNLP, ECCV/ICCV, etc.
- Experience designing rigorous evaluation methodology for generative or agentic systems, including human evaluation and automated benchmarks.
- Experience taking research prototypes towards production, and fluency in one or more of Python, C# or TypeScript.
Please note: Applications are expected to remain open until early September 2026. We encourage candidates to apply as soon as possible, as applications will be reviewed on a rolling basis.
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.
Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.