Skip to main content
Apply Now

Data Science Intern, Asset Intelligence

Job ID REQ-058421
Location Albany, NY
Job Type Intern/Apprentice (Trainee)
Apply Now

The Asset Intelligence Service is the layer that turns a customer's messy asset list into a standard record. Instruments arrive as free text, with the manufacturer misspelled several different ways, two model numbers that differ only by a dash, and a description scraped from a page that was selling something. The service normalizes that text, classifies the instrument, and enriches the record with the attributes the rest of the platform needs.

Almost everything downstream groups on that classification. Fleet and benchmark reports, capital planning, relocation and install planning, and the attribute schema that says what a laboratory instrument of a given class should even have. When a classification is wrong, it is wrong in every one of those places at once. That is why accuracy here earns more attention than a data-cleaning job usually gets.

As Data Science Intern for Asset Intelligence, you learn the service properly and then improve it on a measured basis. The first stretch is learning: trace records end to end, reproduce the pipeline on a known set, and find out where it makes its decisions and on what evidence. After that you own a defined improvement workstream, with a number attached to it that you are responsible for moving.

The role is deliberately narrow. The team is small and the work is live, so ten hours a week goes furthest on one measurable thing owned properly.

Job Responsibilities

Learning the Service

  • Trace an asset record from raw customer intake through normalization, classification, and enrichment to the form it takes in the platform, and be able to explain each step to someone who has never seen it.
  • Reproduce the pipeline on a sample set and confirm you get the same answers the production service does.
  • Write down where the service makes a judgment call, what evidence it uses, and which of those calls are the fragile ones. This document does not exist today and the team needs it.

Measurement

  • Build and maintain a gold-standard reference set of correctly classified instruments, which is the thing the service is currently missing and the thing every accuracy claim depends on.
  • Score match rate and classification accuracy against that set, reported by equipment class rather than as one headline number, so the weak classes are visible instead of averaged away.
  • Re-score after every change, so improvement is demonstrated rather than asserted.

Improvement Work

  • Cluster near-duplicate records and reconcile them as a group with their alias set, rather than one record at a time. Records that differ only by punctuation or spacing are the case the current one-at-a-time approach cannot resolve.
  • Work on the text that feeds the classifier, separating technical capability language from application and use-case language. Scraped marketing copy about what an instrument has been used for is a known source of misclassification.
  • Analyze attribute coverage across the record base, field by field, so the team knows which attributes are populated well enough to build on and which are not before anything gets promised to a customer.

Documentation and Handoff

  • Keep the code and queries in the team's repository (notebooks are fine) in a state someone else can pick up and run.
  • Write short method notes alongside the code. Much of the value of this role is in the write-up, because the point is that the team can repeat and extend the work after the term ends.

Shared Team Work

  • Take on ad hoc data pulls and analysis the team needs at short notice. Everyone on a team this size carries some of that, and it is also the fastest way to see how the platform is actually used.

Enterprise Impact

Under this role, asset classification shifts from:

  • Accuracy inferred from spot checks → accuracy measured against a maintained reference set
  • Records reconciled one at a time → near-duplicates reconciled together with their aliases
  • Errors that spread on their own, because each new record is weighted against how similar ones were already classified → errors found and corrected at the source
  • Method held in one person's working knowledge → method documented and repeatable

Business Outcomes

  • A maintained gold-standard set and a published accuracy baseline the team can hold itself to
  • Measured improvement in classification accuracy, concentrated on the equipment classes where it is currently weakest
  • Attribute coverage visible field by field, so downstream module gaps are known before they are committed to
  • The work documented well enough that it continues after the term ends

Nothing in this job description restricts management's right to assign or reassign duties and responsibilities of this job at any time.

Basic Qualifications

  • Currently enrolled in a bachelor's degree program in data science, statistics, computer science, or a related field.
  • Working knowledge of Python and SQL.
  • Available approximately ten hours per week during the academic term.

Preferred Qualifications

  • Coursework or project experience in data cleaning, record matching, deduplication, or classification.
  • Familiarity with pandas and Jupyter, or the equivalent in R.
  • Some exposure to AI-assisted text processing, along with the instinct to verify what it returns rather than accept it. Knowing how a plausible wrong answer happens matters more here than knowing how to prompt.
  • Comfortable working remotely from written direction with limited supervision, and inclined to ask early when something is ambiguous instead of guessing and proceeding.
  • Clear written communication. The analysis is only worth what the write-up conveys.
  • Interest in scientific instrumentation or laboratory operations. No prior domain knowledge is assumed and none is required to start.

The hourly compensation range for this position is $(25) to $(40) per hour.

Apply Now

Sign Up For Job Alerts

Don’t see what you’re looking for?

Sign up and we’ll notify you when roles become available.

Interested In

By submitting your info, you acknowledge that you have read our privacy policy (this content opens in new window) and consent to receive email communication form.