Responsible AI:

Getting Set Up — Environments,

Tools,

and Avoiding Lock-In

Discuss evaluation as a method for deliberate improvement

Consider aspects of harness engineering

Share possible harnesses

Move from the browser to an agentic coding environment

Goals

Personal

OpenAI 

Anthropic 

Google 

and other commercial LLM providers

Enterprise

AI Sandbox (Portkey)

GitHub Copilot

Claude PU Enterprise (faculty sponsor, staff)

 

see: https://dais.princeton.edu/resources

 

Local (on your computer or HPC)

Anything LLM

Ollama

LM Studio

MLX (Apple Silicon)

vLLM

Before you identify which models are best for a given project, consider your requirements for cost, reproducibility, confidentiality, and compute capacity.   

AI Model Access Options 

Confidential Data

Confidential Data

Confidential Data

Agentic coding assistants like Claude Code or Codex add the ability to work with a folder of files on your computer.

  • Specification files describe, in text, the project goals and requirements.
  • Project data files
  • Scripts to transform and process the data, content, and media.
  • Documentation that explains to others what you're doing.
  • Tests and evaluation that confirm we've actually done what we said we'd do.

The Harness

 

  • VSCode
    • Coding IDE (Integrated Development Environment)
    • Files, code editor, terminal, and chat all in one window
    • Chat is GitHub Copilot by default
    • Plugins for Claude Code, Ollama
  • Zed
    • Zed hosted models by default
    • Directly edit files or chat
  • Claude Code, OpenAI Codex
    • Chat-forward interface with access to files (artifacts)
    • vendor lock-in

Choosing a Harness

I code

It codes

Spec.md - reusable reference file about the larger goals, guidance, and requirements of your project. Usually written with/by the LLM. 

 

Skills.md - Reusable description of a capability or resource

 

MCP and WebMCP (model context protocol) - An external resource or tool. For example tools to search Princeton's library catalog

 

Ponytail - A package of software engineering best practices and opinions

Fitting your harness

Pre-loaded MCP and skills for scientific literature and evidence synthesis (primarily life sciences).

 

Designed for computational analysis using a local machine and high-performance computing

 

Designed for reproducibility. All "artifacts" have a history and logging.

Example:

Claude Science (beta)

  • Domain-tuned version of the Astra model (gpt-6-astra-law)
  • Legal search index — a purpose-built retrieval tool covering U.S. case law, statutes, regulations, court rules, and administrative decisions
  • Custom skills for legal analysis and writing 
  • 47 law-related plugins (LegalQuants, LECG, and Skills.law)

"Evals are the feedback mechanism that lets a harness learn, self-correct, and improve."

Antaripa Saha, Hamel Husain, and Hugo Bowne-Anderson

Don't just test results. Evaluate your inputs, transformations and outputs. 

Establish known inputs, transformations, and outputs.

While the harness turns capability into work, it also adds complexity. How do we know that the model chose the right tool? That the tool-call led to improvement? Where did it add new errors? Evaluation is the craft of making those processes explicit and managable.

From the very beginning, in your specifications, call for code tests and evaluation.

 

An Eval-Centric Harness

  • Research projects are a long-haul. AI can do serious work with the right setup.
  • Identify the right model infrastructure 
  • Find the right harness
  • Configure your harness with relevant tools, resources, and skills. If you need to look it up, let the model look it up (ground everything).
  • Plan for ongoing evaluation of your team. Keep them fed, warm and hydrated. They'll carry you where you need to go.

 

Conclusions