Expoint - all jobs in one place

Finding the best job has never been easier

Limitless High-tech career opportunities - Expoint

Microsoft Research Intern - Training-Time Provenance Data Dignity 
United States, California, Mountain View 
963536968

10.12.2024

Training-time provenanceis a research effort on estimating the influence of specific training data on outputs of large language models (LLMs). Current neural network architectures are opaque in terms of providing sources for their generations, and there are at least two good reasons to change this:

  1. “X-ray” into intent, so that we can detect bad human actors or dangerous AI activity by identifying the most influential source documents related to a given model output. For instance, sneaky prompts might invoke articles about bomb making that could evade guardrails otherwise. This will be a deeper method of countering this type of danger than others currently in use.
  2. “Data dignity”, meaning incentives, recognition, and potentially pay for people who contribute certain valuable data to unforeseen kinds of models we will want in the future, assuming the future will surprise us fundamentally. The goal is to foster new classes of creative professionals where possible, instead of relying solely on ideas like Universal Basic Income in the event of a future with very high-functioning large models.
Required Qualifications
  • Currently enrolled in a PhD program in Computer Science or a related STEM field. Exceptional candidates enrolled in a master’s program might also be considered.
  • At have at least 2 years of research experience, including peer-reviewed publications, researching a topic closely related to the above description, such as natural language processing, deep learning, generative models, approximation methods, etc.

Other Requirements

  • Research Interns are expected to be physically located in their manager’s Microsoft worksite location for the duration of their internship.
  • In addition to the qualifications below, you’ll need to submit a minimum of two reference letters for this position as well as a cover letter and any relevant work or research samples. After you submit your application, a request for letters may be sent to your list of references on your behalf. Note that reference letters cannot be requested until after you have submitted your application, and furthermore, that they might not be automatically requested for all candidates. You may wish to alert your letter writers in advance, so they will be ready to submit your letter.
Preferred Qualifications
  • Demonstrated ability to develop original research agendas.
  • Ability to collaborate effectively with other researchers and product development teams.
  • Experience in training large AI models.
  • Experience in approximation methods for deep learning systems.
  • Proficient interpersonal skills, cross-group, and cross-culture collaboration.
  • Ability to think unconventionally to derive creative and innovative solutions.

Applied Sciences IC2 : The base pay range for this internship is USD $5,460 -$10,680 per month.Applied Sciences IC3 : The base pay range for this internship is USD $6,550 -$12,880 per month.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:

Responsibilities

For this Research Internship (summer 2025), we are seeking PhD students with a passion for fundamental Deep Learning research, particularly those with experience in training LLMs and other large AI models. The Research Intern's responsibilities will include (1) training small language models with novel schemes preserving provenance of data, (2) experimenting with these models to test their performance and reliability.