Something about me

What I do, what I work with, and recent things I have built.

James Conner

James Conner
Infrastructure Operations Manager, Institute for Health Metrics and Evaluation, University of Washington

I run infrastructure and data center operations at IHME, part of the University of Washington. Before that I spent eight and a half years in data center operations for VMware and Amazon Web Services, and I started out as a network technician in the Army in 2007.

Most of that career has been the same job wearing different badges. Keep critical systems available, know why they broke, and build teams and procedures that hold up when something large fails at an inconvenient hour. The parts I care most about are incident response, making failure visible before a customer finds it, and getting operational knowledge out of people's heads and into something the next person can read.

I am happy where I am and not looking to leave, but I am always open to a good conversation about operations, automation, or infrastructure work.

What I work with · Recent work · Experience · Education · Elsewhere

What I work with

At data center scale, professionally. Data center operations management, incident command for large-scale events, capacity and power planning, disaster recovery and backup policy, change and configuration management, SLA and vendor management, budget ownership, and hiring, mentoring and performance managing operations teams.

Hands on and current. Linux, Docker and Compose, Proxmox virtualization and clustering, Ansible, pfSense with VLAN segmentation, reverse proxy with single sign-on, internal certificate authority and PKI, secrets management with SOPS and age, Grafana, InfluxDB, Loki and Telegraf, CI with Gitea Actions, GitOps, n8n workflow automation, and local LLM inference on GPU.

I split those two lists on purpose. The first is backed by payroll and headcount. The second is backed by a lab I run at home, where I get to make the mistakes somewhere they cost nobody anything.

Recent work

That lab is built to be operated the way a real environment is operated, rather than assembled and abandoned. Everything below is dated, because infrastructure writing goes stale and a reader deserves to know how old a claim is.

Self-healing metrics collection (August 2026). Monitoring that fails quietly is worse than no monitoring, because it reports healthy right up until you need it. I built automatic recovery for every pull-based metrics collector in the fleet and proved it against real induced faults rather than rehearsals. Fault to alert to fresh data landed inside 90 seconds in both trials. The success criterion is what made it work. The check is whether a new data point actually arrived in the database, never whether the container claims to be running, because that second question was the one that had been answering wrong all along.

Writing it up (August 2026). Published Sync your Spotify track to your Slack status with n8n, a build along for people who have never used the tool. The expensive part turned out to be verification rather than composition. One widely repeated API limit was wrong in every write up I could find, including my own memory of it, and only fell over when I read the vendor's own current documentation.

Continuous integration and guardrails (July 2026). A CI runner that gates every change to the infrastructure repository on secret scanning across full history, shell and YAML linting, and a documentation build that fails on a broken link. Both gates were tested in both directions, since a check that has never been seen to fail is not yet a check.

Earlier

WhenWhatResult
July 2026Host and hardware metricsEvery host, plus patch status, hypervisor storage and GPU health. Caught a GPU silently falling back to CPU.
July 2026Fleet-wide log aggregationSearchable forensic record across the fleet. Deliberately not an alert source.

Experience

University of Washington · Infrastructure Operations Manager · 2024 to present

I lead the team that runs IHME's data center, covering day-to-day operations, monitoring infrastructure and conference room technology. I set disaster recovery and backup procedures, own the budget, and write the policies that protect IHME's IT assets and data. I supervise staff and provide technical leadership to the wider IT team, work with vendors and contractors on cost-effective service delivery, and plan system upgrades and rollouts to keep downtime low.

VMware · Data Center Operations Manager · 2019 to 2024

Data Center Operations Manager
February 2023 to January 2024 (1 year)

I was operationally responsible for one of VMware's largest data centers, running day-to-day operations while developing longer-term strategy. I managed people across a wide range of skill sets, mentored them, and gave direct performance feedback. I sponsored projects, managed work queues, and proposed technical solutions as issues came up. I prepared for and responded to large-scale events, and produced root cause analysis to prevent repeats. I recommended, documented and oversaw policies and procedures against industry best practice, which is what let the site exceed its service level agreements. I reported weekly and quarterly to senior data center management.

Data Center Operations Supervisor
September 2019 to February 2023 (3 years 6 months)

I ran all aspects of data center operations, including maintaining and repairing critical IT and R&D equipment to keep supported services available. I supervised the local operations team and made sure remote hands work, customer installations and service requests were done on time. I planned and coordinated ongoing deployments and modifications against current and future demand. As manager for the technical staff handling most break/fix incidents across business units, I drove global support for local site ticket resolution.

I scheduled on-site and on-call staff, and owned inventory accuracy and configuration records for deployed solutions. I delivered weekly ticket management reporting covering SLA results and ticket aging. I took part in incident response, resolved escalated service issues, and developed process improvements for the customer experience. Working with other teams I shortened deployment schedules and verified deployment detail, and I reviewed power distribution and scheduled modifications for power balancing at the systems level.

Amazon Web Services · Data Center Operations Manager · 2015 to 2019

Data Center Operations Manager
August 2017 to September 2019 (2 years 2 months)

I managed and mentored a team of diverse professionals, setting strategy for resolving customer technical issues while supporting the team's own growth. I performance-managed individuals against the expectations of their roles. I recommended, documented and oversaw policies and procedures against industry best practice, which let us exceed required service level agreements. I reported weekly and quarterly to senior data center management on the key performance indicators affecting fleet health. As overall manager for servers and networks in the building I owned compliance with security and safety policy. During large-scale outages I was call leader for multiple technical teams through recovery.

Data Center Tech Lead
November 2016 to August 2017 (10 months)

I led a team of technicians and guided how we implemented and improved data center operating processes. I trained new personnel and ran ongoing training for experienced staff as new technology arrived. I developed and implemented policies and procedures aimed at team efficiency and productivity. I installed and maintained physical data center infrastructure, including server administration, repair and hardware troubleshooting in a Linux environment.

Data Center Technician II
July 2015 to November 2016 (1 year 5 months)

I used diagnostic tools to identify and fix data center hardware faults with minimal downtime to critical systems. Working under minimal supervision, I balanced competing priorities through the ticketing system and worked with internal and external teams on complex customer issues. I provided operational support to data centers in the Oregon area, covering day-to-day incident management of servers and networking equipment along with project work and capacity management. I took part in the on-call rotation for after-hours support, and I wrote and updated standard operating procedures and knowledge articles to speed up ticket resolution.

US Army · Information Systems Manager and Network Specialist · 2007 to 2008, 2013 to 2015

Information Systems Manager / Network Specialist
August 2014 to July 2015 (1 year)

  • Configured and managed multi-vendor information processing equipment in the operating configurations required for mission support.
  • Maintained and troubleshot more than 500 classified and unclassified computer systems across multiple networks, and managed over $500,000 in COMSEC and automation equipment.
  • Worked with cross-functional teams to develop and implement custom firewalls and applications able to withstand advanced cyberattacks, protecting critical logistical and data systems for the Army.
  • Ran training sessions so team members could use the networks and software tools effectively.
  • Organized and facilitated team building activities, and took part in Army professional development programs and community engagement events.

Network / Junior System Administrator
August 2013 to July 2014 (1 year)

  • Supervised and mentored a team of 5 IT specialists on triage and troubleshooting of Tier 2 and Tier 3 issues.
  • Worked with organization customers daily to resolve IT issues and keep the unresolved ticket count low.
  • Managed the company's shared drive servers for availability and correct customer access.
  • Took part in information systems and infrastructure projects that improved network availability, security and routing, maintaining 99.98% uptime on informational assets.

Tier II Network Technician
July 2007 to July 2008 (1 year 1 month)

  • Led a team of network technicians configuring and optimizing hardware and software for over 500 employees.
  • Worked with cross-functional teams to implement IT solutions against customer needs while holding 99.98% network uptime.
  • Troubleshot and resolved complex switch and connectivity faults, leading the team on root cause and fixes.
  • Supervised and managed $500,000 of highly sensitive equipment with zero loss or breach over a one-year period.
  • Mentored and coached junior team members.

Education

Capella University
Bachelor of Science, Information Technology (2015 to 2020)

Washington State University
Master of Business Administration (in progress)

Certifications

Both have lapsed. They are listed for completeness rather than as current credentials.

Cisco Certified Network Associate (CCNA), Cisco. Issued March 2015, expired March 2018.
Security+, CompTIA. Issued November 2014, expired November 2017.

Elsewhere

LinkedIn