
Nicholas Steinwachs
Jul 21, 2026
Looking for a principal-level devops / infrastructure expert to help bring us from good to great while guiding cost-control and performance improvements on a fractional basis.
Team: Maintenance & DevOps
Engagement: Independent contractor, ongoing, approximately 4-8 hours per week (roughly 10-20% of full time)
Location: Remote, East Coast (US)
Works with: Nick Steinwachs (President), Molly Steinwachs (COO), and Max Kadel (Sr. DevOps), who leads the platform; provides review and support to our Mid-Level DevOps / Infrastructure Engineer
Compensation: $200 - $250/hour (target $225/hour); major-incident availability beyond base hours structured separately
About Notch8
Notch8 is a family-owned software company and Samvera Partner since 2016. We build and host open-source digital repository and digital collections platforms for academic libraries, consortia, research institutions, and cultural-heritage organizations. We run Hyku, Manifold, InvenioRDM, Dataverse, and supporting applications for customers including PALNI/PALCI, the University of Tennessee, West Virginia University, Princeton, and many other leading institutions and consortia. We already work with a number of contractors, so a fractional senior engagement fits how we operate.
The engagement
We host a multi-tenant, multi-application platform on AWS, and we are growing it on three fronts: onboarding new repository platforms onto shared infrastructure, bringing our hosting cost structure down substantially, and raising reliability as we take on larger customers. At the same time, we are working toward SOC 2 Type II, which needs senior ownership of the control environment.
Our senior engineer leads the platform and owns day-to-day operations and direction. We are looking for a principal-level engineer to add depth on a fractional basis: to own the technical side of our SOC 2 program, provide principal-grade backup and review on our highest-risk architecture and security decisions, and help level up the DevOps team. You will align with our platform lead on strategy and goals first, then apply your hours where principal experience matters most. This is a high-leverage, mostly asynchronous engagement for someone who can move a platform forward without needing to touch everything themselves.
What you will own
Technical ownership of the SOC 2 Type II program: control design, the evidence strategy that makes control operation continuous and auditable, liaison with our auditor, and direction on gap remediation. Type II requires that controls demonstrably operate over a sustained period, and you will own getting us there and keeping us there.
Principal-grade review and backup on the highest-risk architecture and security decisions for the AWS and Kubernetes platform, including the roadmap for onboarding new hosted applications onto shared infrastructure, working alongside the platform lead who owns that direction.
Advisory and review on the hosting cost-reduction roadmap: right-sizing and autoscaling for bin-packing, Spot adoption for suitable workloads, storage tiering and migration, cluster consolidation, and commitment-based discounts. We have a tiered plan targeting a reduction of half or more, and your judgment sharpens it.
Review of reliability standards: service-level objectives, highly-available and backed-up data stores, tested disaster-recovery procedures, and a maturing incident-response practice.
Review and support for the Mid-Level engineer, helping level them toward Senior, and a technical sounding board for the platform lead.
Senior escalation for major incidents, under a defined availability arrangement beyond the base weekly hours.
Our stack
AWS (EKS, EC2, EFS, S3, RDS, networking), us-west-2.
Kubernetes, Helm, and ArgoCD for GitOps-based deployment.
OpenTofu for infrastructure-as-code.
PostgreSQL via Kubernetes operators with operator-managed backups; Redis and Solr as supporting services.
Prometheus and Grafana for metrics; cost-allocation tooling for spend visibility.
1Password with externalized secret references; Cloudflare and nginx-ingress at the edge.
Ruby on Rails applications (Hyku, Manifold, and others) running on the platform.
What we are looking for
A principal- or staff-level track record operating production infrastructure at scale on AWS, with deep Kubernetes (EKS) experience.
Fluency with infrastructure-as-code (OpenTofu or Terraform) and GitOps deployment (ArgoCD, Flux, or similar).
Hands-on SOC 2 Type II implementation experience, or equivalent depth with ISO 27001, FedRAMP, or HIPAA control environments.
A demonstrated record of measurably reducing cloud cost through architecture and commitment strategy.
Experience with multi-tenant SaaS or platform infrastructure serving multiple customers from shared resources.
The ability to work at high leverage on limited hours: strong asynchronous communication, sharp prioritization, and the judgment to know which decisions genuinely need principal attention.
Comfort operating as backup, reviewer, and mentor, with day-to-day execution and platform direction owned by the internal team.
Nice to have
Experience with compliance-automation platforms (Vanta, Drata, Secureframe, or similar).
Karpenter, Spot orchestration, or EKS cost optimization at scale.
PostgreSQL operated at scale via Kubernetes operators (Zalando, CrunchyData, or comparable).
Familiarity with the Samvera / Hyku ecosystem or with academic-library and research-data platforms.
Engagement terms
Independent contractor, remote, ongoing with quarterly review of scope and hours. Base commitment of roughly 4-8 hours per week at $200 - $250/hour (target $225/hour), with a defined arrangement for additional availability during major incidents billed separately (retainer or premium multiple).