Responsible AI:
Getting Set Up — Environments,
Tools,
and Avoiding Lock-In
Discuss evaluation as a method for deliberate improvement
Consider aspects of harness engineering
Share possible harnesses
Move from the browser to an agentic coding environment
Goals
Personal
OpenAI
Anthropic
and other commercial LLM providers
Enterprise
AI Sandbox (Portkey)
GitHub Copilot
Claude PU Enterprise (faculty sponsor, staff)
see: https://dais.princeton.edu/resources
Local (on your computer or HPC)
Anything LLM
Ollama
LM Studio
MLX (Apple Silicon)
vLLM
Before you identify which models are best for a given project, consider your requirements for cost, reproducibility, confidentiality, and compute capacity.
AI Model Access Options
Confidential Data
Confidential Data
Confidential Data
Agentic coding assistants like Claude Code or Codex add the ability to work with a folder of files on your computer.
- Specification files describe, in text, the project goals and requirements.
- Project data files
- Scripts to transform and process the data, content, and media.
- Documentation that explains to others what you're doing.
- Tests and evaluation that confirm we've actually done what we said we'd do.
The Harness
- VSCode
- Coding IDE (Integrated Development Environment)
- Files, code editor, terminal, and chat all in one window
- Chat is GitHub Copilot by default
- Plugins for Claude Code, Ollama
- Zed
- Zed hosted models by default
- Directly edit files or chat
- Claude Code, OpenAI Codex
- Chat-forward interface with access to files (artifacts)
- vendor lock-in
Choosing a Harness
I code
It codes
Spec.md - reusable reference file about the larger goals, guidance, and requirements of your project. Usually written with/by the LLM.
Skills.md - Reusable description of a capability or resource
MCP and WebMCP (model context protocol) - An external resource or tool. For example tools to search Princeton's library catalog
Ponytail - A package of software engineering best practices and opinions
Fitting your harness
Pre-loaded MCP and skills for scientific literature and evidence synthesis (primarily life sciences).
Designed for computational analysis using a local machine and high-performance computing
Designed for reproducibility. All "artifacts" have a history and logging.
Example:
Claude Science (beta)
- Domain-tuned version of the Astra model (gpt-6-astra-law)
- Legal search index — a purpose-built retrieval tool covering U.S. case law, statutes, regulations, court rules, and administrative decisions
- Custom skills for legal analysis and writing
- 47 law-related plugins (LegalQuants, LECG, and Skills.law)

"Evals are the feedback mechanism that lets a harness learn, self-correct, and improve."
Antaripa Saha, Hamel Husain, and Hugo Bowne-Anderson
Don't just test results. Evaluate your inputs, transformations and outputs.
Establish known inputs, transformations, and outputs.
While the harness turns capability into work, it also adds complexity. How do we know that the model chose the right tool? That the tool-call led to improvement? Where did it add new errors? Evaluation is the craft of making those processes explicit and managable.
From the very beginning, in your specifications, call for code tests and evaluation.
An Eval-Centric Harness

- Research projects are a long-haul. AI can do serious work with the right setup.
- Identify the right model infrastructure
- Find the right harness
- Configure your harness with relevant tools, resources, and skills. If you need to look it up, let the model look it up (ground everything).
- Plan for ongoing evaluation of your team. Keep them fed, warm and hydrated. They'll carry you where you need to go.
Conclusions
deck
By Andrew Janco
deck
- 95