Full-Time

Senior Cloud Architect

Delivery, GenAI

DoiT International

DoiT International

501-1,000 employees

Cloud consulting and optimization platform

No salary listed

Remote in Canada

Remote

Remote from Canada; EST preferred.

Bachelor's

Category
DevOps & Infrastructure (1)
Required Skills
Redshift
TensorFlow
PyTorch
Docker
AWS
Terraform
Google Cloud Platform

Get referred to DoiT International

See people who can refer or advise you

Requirements
  • 4+ years of experience architecting, deploying, and managing cloud-based AI/ML solutions, including production workloads.
  • Proven track record designing and operating large, distributed systems on AWS, selecting appropriate services and patterns to meet business and technical goals.
  • Advanced proficiency with AWS services relevant to AI/ML and GenAI.
  • Hands-on experience with Amazon Bedrock for deploying and scaling foundation models and Generative AI workloads.
  • Experience fine-tuning and deploying Large Language Models and multimodal AI using Amazon SageMaker (including JumpStart).
  • Strong prompt engineering skills and familiarity with rigorous model evaluation (quality, safety, performance).
  • Understanding of agentic capabilities and patterns for AI agents that autonomously perform tasks and integrate with existing systems.
  • Experience with Amazon Q Business and Amazon Q Developer (or similar tools) to accelerate insight generation and development workflows.
  • In-depth knowledge of Amazon SageMaker components such as Pipelines, Model Monitor, Data Wrangler, and SageMaker Clarify for bias detection and interpretability.
  • Proficiency integrating TensorFlow, PyTorch, and other ML frameworks with SageMaker for model development, fine-tuning, and deployment.
  • Experience with distributed training (multi-GPU or multi-node) and performance optimization for inference.
  • Strong data-engineering skills on AWS: Amazon S3, AWS Glue, Lake Formation, Redshift for AI/ML data pipelines.
  • Experience building end-to-end AI/ML workflows using services like AWS Lambda, Step Functions, API Gateway, and containerized deployments on Amazon EKS / AWS Fargate.
  • Hands-on experience with CI/CD for AI/ML using AWS CodePipeline, CodeBuild, SageMaker Pipelines, or similar.
  • Proficiency in monitoring and operating AI systems using Amazon CloudWatch and SageMaker Model Monitor.
  • Strong understanding of AI governance, security, and compliance on AWS, including IAM, KMS, and data privacy patterns.
  • Familiarity with AI ethics and bias detection/mitigation (e.g., using SageMaker Clarify or similar tools).
  • Working knowledge of Google Cloud AI tools sufficient to reason about multi-cloud architectures and integration points.
  • Proven ability to mentor peers, run enablement sessions, and collaborate across Sales, CS, and Product.
  • Excellent communication skills across technical and business audiences; able to simplify complex ideas and influence decisions.
  • Natural ownership mentality: you escalate early, resolve fast, and own the outcome.
  • Demonstrated ability to work effectively in a remote-first, global environment.
Responsibilities
  • Be the trusted cloud engineer customers lean on for high‑impact technical optimization work across cost, reliability, security, and performance.
  • Design and help implement solutions that improve cost efficiency (rightsizing, reservations/commitments, storage optimization, etc.).
  • Design and help implement solutions that increase reliability and resilience (HA/DR architectures, SLO/SLA‑aware designs).
  • Design and help implement solutions that strengthen security posture (IAM, network segmentation, data protection, least‑privilege).
  • Design and help implement solutions that reduce operational toil (automation, self‑service, guardrails, policy enforcement).
  • Plan and deliver structured engagements such as Cloud Optimization Sessions, cost/efficiency/performance workshops, security posture or reliability reviews, and architecture deep dives / well‑architected style assessments.
  • Respond to Expert Inquiry / support requests that require deep cloud engineering expertise, ensuring high‑quality, well‑explained resolutions.
  • Bring domain depth in ML / GenAI – deploying and operating ML/GenAI workloads (training and inference), GPU utilization, scaling, and cost control; MLOPS and integrating workloads with monitoring, logging, and FinOps; safe and efficient use of managed AI services.
  • Convert one‑off customer solutions into Gravel Roads - reusable patterns such as playbooks, Terraform modules, CloudFlow templates, cloud diagrams, Composer Recipes -> DCI Insights, and internal /external documentation.
  • Provide structured feedback to the DoiT Product and Engineering teams on product gaps and friction points discovered in real‑world usage, new opportunities for automation and workload lenses within DCI, telemetry and tracking that would make future FDE work more efficient.
  • Contribute directly to DCI where appropriate - from feature requests and feedback, to contributing code, to owning specific DCI features end‑to‑end.
  • Build agent skills, scripts, and internal tooling that codify your expertise and scale it across the team.
  • Contribute to internal enablement: share learnings via documentation, demos, office hours, or training sessions for other FDEs and Customer Success team members.
  • Operate as an embedded technical partner inside the account team.
  • Work in the account team model alongside Customer Success Managers and Account Managers to deliver impactful outcomes.
  • Own the technical depth lane: technical deployment & integration, automation & platform adoption, signal‑based proactive engagement, and repeatable Cloud Optimization solutions.
  • Partner with customers' engineers, architects, and FinOps teams to translate vague pain points into concrete technical optimization plans and help them ship changes that stick and create continuous value.
  • Co‑deliver complex or multi‑domain engagements with peer FDEs (infra + data + ML/GenAI), reviewing and refining designs, engagement plans together.
  • Communicate complex technical topics clearly to engineers and non‑technical stakeholders, and maintain clear documentation of architectures, decisions, and implemented changes.
  • Contribute to a culture of continuous improvement within the global FDE community through design reviews, internal forums, enablement sessions, and experimentation.
  • Become an expert in DoiT Cloud Intelligence products and services including Cloud Analytics, DCI Insights, Cloud Composer, CloudFlow, DataHub, PerfectScale, and other Enterprise Platforms.
  • Use DCI hands-on to build and operationalize Cloud Analytics and Allocations to create dashboards and reports for customer engineering, finance, and leadership.
  • Use DCI Insights to identify and prioritize cost, risk, and reliability opportunities, and shepherd them through to closure.
  • Implement Cloud Composer queries, build recipes that result in hand-crafted insights across all customers' engineering use cases.
  • Build CloudFlow automations (e.g., anomaly routing, scheduled actions, guardrails, policy enforcement).
  • Use Built in Integrations such and utilize DataHub and other workload‑intelligence features to optimize key business and workload data inside DCI.
  • Help customers embed DCI into existing observability, CI/CD, and governance processes so it becomes trusted and indispensable in day‑to‑day cloud operations.
Desired Qualifications
  • BA/BS degree in Computer Science, Mathematics, or a related technical field, or equivalent practical experience.
  • Additional data or AI certifications (e.g., AWS/GCP data certifications, reputable AI/ML programs such as Stanford, Coursera, Udacity, MIT, eCornell).
  • Experience with modern RLHF, advanced fine-tuning techniques, and hybrid AI architectures.
  • Familiarity with Hugging Face or similar open-source ecosystems integrated with AWS.
  • Prior experience as a ML Engineer, Data Scientist, or AI-focused Architect in a consulting or SaaS environment.
  • Experience with JIRA or similar tools for tracking work across delivery and product-feedback cycles.
  • Exposure to Agile practices and frameworks commonly used for SaaS and cloud delivery.

DoiT International provides cloud consulting and technology solutions to help businesses optimize their cloud operations. The company offers expert services alongside proprietary software tools to manage cloud environments, focusing on cloud migration, cost optimization, and performance tuning. It collaborates with major cloud platforms such as Google Cloud and AWS to deliver tailored solutions and subscriptions, generating revenue from consulting fees and software licenses. DoiT differentiates itself through its direct partnerships with large cloud providers, its suite of in-house tools, and a blended business model that combines one-time consulting with recurring software subscriptions. Its goal is to help clients accelerate digital transformation, maximize cloud investments, and maintain efficient, scalable cloud environments for a diverse client base.

Company Size

501-1,000

Company Stage

Series A

Total Funding

$100M

Headquarters

San Francisco, California

Founded

2011

Get referred to DoiT International

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • May 2026 PerfectScale for Commitments targets dynamic Kubernetes and autoscaling waste immediately.
  • June 2026 Ingram Micro alliance exposes DoiT to 5,000-plus AWS partners.
  • July 2026 Attribute launch attacks AI token spend, a fast-growing budgeting pain point.

What critics are saying

  • AWS Cost Explorer remains free and native, pressuring DoiT on AWS-only accounts.
  • DoiT’s acquisition spree, including SELECT and Attribute, raises integration and execution risk.
  • If AWS or GCP bundles deeper FinOps automation, DoiT loses its reseller-led wedge.

What makes DoiT International unique

  • DoiT’s June 2026 AWS BVR Competency proves measurable business-outcome delivery, not just tooling.
  • PerfectScale uses workload-aware continuous optimization across Kubernetes, commitments, Snowflake, and Databricks.
  • DoiT’s forward-deployed engineers and 99.7% satisfaction differentiate it from software-only FinOps vendors.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Unlimited Paid Time Off

Flexible Work Hours

Health Insurance

Parental Leave

Employee Stock Option Plan

Home Office Stipend

Professional Development Budget

Peer Recognition Program

Company News

Precedence Research
Jul 28th, 2026
DoiT acquires Attribute to launch real-time AI tokenomics and earns AWS Business Value Realization status

DoiT acquired AI cost-management startup Attribute in July 2026 to launch a real-time, kernel-level tokenomics product. The Santa Clara-based company has acquired five companies in the past 18 months, focusing on AI cost management. DoiT simultaneously became an inaugural partner for the new AWS Business Value Realization Competency, which proves measurable financial and technical outcomes for enterprise customers. This recognition complements a five-year strategic collaboration agreement with AWS to generate $5 billion in business. The acquisition enables DoiT to offer automated, zero-instrumentation cost attribution that tracks real-time consumption at the kernel level. The platform connects AI token costs directly to individual customers, features, and autonomous agents without requiring SDKs or tagging policies. DoiT aims to lead in establishing operational standards for AI token attribution alongside the newly formed Tokenomics Foundation by the Linux Foundation and FinOps Foundation.

PR Newswire
Jul 7th, 2026
DoiT launches Attribute to track AI spend without code changes

DoiT has launched Attribute™, a technology that tracks AI spending across tokens, model requests, and GPU usage without requiring SDKs, tags, or code changes. The system uses a lightweight eBPF sensor that installs in approximately 15 minutes and immediately provides detailed cost breakdowns per customer, feature, and agent. The technology addresses a growing challenge in AI cost management. DoiT projects monthly AI spend will triple in the next 12 months, yet only 15% of 500 surveyed enterprise leaders said they could calculate AI ROI without significant bottlenecks. Attribute™ measures consumption at the operating system kernel level, tracking every GPU cycle, API call, and token back to the specific workload, tenant, or agent responsible. The system automatically splits cached, reasoning, input, and output tokens and integrates with cost data from Anthropic, OpenAI, Google Gemini, and AWS Bedrock.

CRN
Jul 7th, 2026
DoiT Buys AI FinOps Startup Attribute, Launches AI Token Cost Management Product

DoiT Buys AI FinOps Startup Attribute, Launches AI Token Cost Management Product

PR Newswire
Jun 16th, 2026
DoiT expands PartnerOps platform with revenue management for cloud distributors

DoiT has expanded its PartnerOps platform with Revenue Management capabilities, responding to growing demand from cloud distributors and hyperscalers. The platform helps channel partners operate profitable cloud practices by automating billing, reconciliation and margin protection across multi-tier distribution structures. The expansion builds on DoiT's November 2025 alliance with Ingram Micro, which brought DoiT Cloud Intelligence to thousands of AWS partners. In December, DoiT earned AWS Managed Services Provider Programme Designation, placing it amongst the top tier of AWS MSP partners globally. Revenue Management automates the conversion of cloud usage into invoice-ready charges, applying contract-aware pricing logic across distributors, resellers and end customers. The module eliminates manual reconciliation work and reduces revenue leakage. DoiT manages over $20 billion in cloud spend for 4,500 customers across 27 countries.

PR Newswire
Jun 9th, 2026
DoiT extends SELECT to Databricks, automating cost optimisation across $250M+ in data platform spend

DoiT has launched SELECT for Databricks, extending its automated cost optimisation platform to help engineering and data teams manage Databricks expenses. The product provides full cost visibility across Databricks environments and underlying cloud infrastructure, combining both into unified reporting. SELECT for Databricks offers automated savings of up to 30% on workloads, granular cost attribution by workspace and team, and machine learning-powered anomaly detection. The platform integrates via read-only access and requires approximately 20 minutes to set up. Built on technology proven across $250 million in Snowflake spend, SELECT now supports Databricks alongside Snowflake, with Google BigQuery in early preview. Clearco reported a 15% usage reduction within two days of deployment. SELECT is available standalone or within DoiT's Cloud Intelligence platform.