Useful Links

Useful Links for Multicloud and AI Architects

Last reviewed: July 3, 2026.

This page is a curated reference library for multicloud architects, hybrid cloud engineers, AI architects, platform engineers, and technical leaders building modern enterprise platforms. It includes official documentation, architecture centers, security frameworks, AI engineering tools, automation references, FinOps resources, observability platforms, and hands-on learning sites.

The goal is simple: keep the most useful architecture, cloud, AI, automation, networking, security, and operations links in one practical place.

Categories

Cloud Architecture and Well-Architected Guidance

These links are useful when designing cloud platforms, reviewing workload quality, planning migrations, creating landing zones, or comparing architecture patterns across providers.

  • Azure Architecture Center
    What it is: Microsoft’s architecture hub for Azure reference architectures, solution ideas, design patterns, and decision guides.
    Why it is useful: Helpful for designing Azure landing zones, hybrid architectures, identity patterns, networking topologies, data platforms, and application modernization strategies.
  • Azure Well-Architected Framework
    What it is: Microsoft’s workload review framework for reliability, security, cost optimization, operational excellence, and performance efficiency.
    Why it is useful: Useful for architecture reviews, design checklists, executive readouts, and identifying technical debt before production deployment.
  • AWS Architecture Center
    What it is: AWS’s architecture guidance hub for reference architectures, diagrams, design patterns, and cloud solution guidance.
    Why it is useful: Helps architects compare AWS-native approaches to Azure, Google Cloud, VMware, and hybrid cloud designs.
  • AWS Well-Architected Framework
    What it is: AWS’s framework for evaluating workloads against key architecture pillars.
    Why it is useful: Great for structured workload reviews, risk scoring, remediation planning, and building repeatable design standards.
  • Google Cloud Architecture Center
    What it is: Google Cloud’s reference architecture and design pattern library.
    Why it is useful: Useful for comparing Google Cloud patterns across networking, data analytics, AI, security, application modernization, and hybrid connectivity.
  • Google Cloud Well-Architected Framework
    What it is: Google Cloud’s framework for designing and operating secure, efficient, reliable, and cost-effective systems.
    Why it is useful: Useful when building design scorecards or comparing workload quality across multiple clouds.
  • Microsoft Cloud Adoption Framework
    What it is: Microsoft’s cloud adoption methodology for strategy, planning, readiness, governance, management, and migration.
    Why it is useful: Helpful for building cloud programs, migration factories, governance models, landing zones, and operating models.
  • AWS Cloud Adoption Framework
    What it is: AWS’s cloud transformation model organized around business, people, governance, platform, security, and operations perspectives.
    Why it is useful: Useful for mapping technology decisions to organizational readiness, staffing, governance, and operating model maturity.

Public Cloud Documentation

These are the primary documentation portals every multicloud architect should keep bookmarked.

  • Microsoft Azure Documentation
    What it is: The main documentation portal for Azure services, architecture, administration, CLI, PowerShell, APIs, and implementation guidance.
    Why it is useful: Essential for validating Azure service behavior, limits, deployment steps, security controls, and operational guidance.
  • AWS Documentation
    What it is: The official documentation portal for AWS services, APIs, SDKs, CLI, architecture, operations, and security.
    Why it is useful: Useful for confirming AWS service configuration, quotas, IAM behavior, network design, automation options, and troubleshooting steps.
  • Google Cloud Documentation
    What it is: The official documentation portal for Google Cloud services, APIs, SDKs, architecture, AI, data, networking, and security.
    Why it is useful: Useful for understanding Google Cloud’s design patterns, service models, IAM, networking, and AI platform capabilities.
  • Azure CLI Documentation
    What it is: Microsoft’s command reference for managing Azure resources from the command line.
    Why it is useful: Valuable for automation, scripting, deployment validation, and troubleshooting in Azure environments.
  • AWS CLI Command Reference
    What it is: AWS’s command reference for managing AWS services through the CLI.
    Why it is useful: Helpful for scripting, diagnostics, automation, and creating repeatable operational workflows.
  • Google Cloud CLI Documentation
    What it is: Google Cloud’s command-line interface documentation.
    Why it is useful: Useful for managing Google Cloud resources, validating deployments, and automating cloud operations.

Hybrid and Private Cloud

These links are useful for architects designing hybrid cloud, private cloud, VMware Cloud Foundation, Azure Local, Nutanix, Dell infrastructure, and enterprise modernization platforms.

  • Broadcom TechDocs
    What it is: Broadcom’s documentation portal for VMware Cloud Foundation, vSphere, NSX, Aria, Tanzu, and related VMware platform products.
    Why it is useful: The primary source for current VMware and VCF documentation after the VMware to Broadcom transition.
  • VMware Cloud Foundation Documentation
    What it is: Official documentation for VMware Cloud Foundation 9.1.
    Why it is useful: Critical for VCF design, deployment, lifecycle management, workload domains, operations, networking, and private cloud architecture.
  • VMware vSphere Documentation
    What it is: Official documentation for VMware vSphere 9.0.
    Why it is useful: Useful for validating ESXi, vCenter, clustering, HA, DRS, storage, networking, lifecycle, and operational behavior.
  • VMware NSX Documentation
    What it is: Official documentation for VMware NSX 4.2.
    Why it is useful: Essential for network virtualization, distributed firewalling, segmentation, overlay networking, Tier-0 and Tier-1 routing, and private cloud security.
  • Azure Local Documentation
    What it is: Microsoft’s documentation for Azure Local, the Arc-enabled distributed infrastructure platform for running virtual machines, containers, and selected Azure services outside Azure regions.
    Why it is useful: Important for architects evaluating Microsoft hybrid cloud, edge modernization, Azure-connected infrastructure, and VMware alternative strategies.
  • Azure Local Baseline Reference Architecture
    What it is: Microsoft’s baseline reference architecture for Azure Local design.
    Why it is useful: Useful for understanding recommended design patterns, management considerations, network planning, and operational baselines.
  • Dell Technologies Info Hub
    What it is: Dell’s technical content hub for infrastructure, storage, cloud, AI, data protection, and workload solutions.
    Why it is useful: Useful for Dell-based architecture references, design guides, validated solutions, and platform-specific engineering detail.
  • Dell PowerFlex Product Documentation
    What it is: Dell’s documentation index for the PowerFlex family.
    Why it is useful: Useful for software-defined storage, private cloud platforms, VCF designs, Kubernetes integrations, and high-performance workload planning.
  • Nutanix Documentation Portal
    What it is: Nutanix’s documentation portal for Prism, AOS, AHV, networking, storage, security, and operations.
    Why it is useful: Useful when comparing VMware, Azure Local, and Nutanix private cloud options.
  • Nutanix Developer Portal
    What it is: Nutanix’s developer site for APIs, automation, SDKs, and infrastructure programmability.
    Why it is useful: Helpful for automation engineers and architects building scripted or API-driven Nutanix workflows.

Networking and Connectivity

These links are useful for designing hybrid connectivity, cloud WAN, segmentation, transit routing, private cloud networking, and multicloud network patterns.

  • Azure Virtual WAN Documentation
    What it is: Microsoft’s documentation for Azure Virtual WAN, hub networking, routing, VPN, ExpressRoute, and secured virtual hubs.
    Why it is useful: Useful for enterprise hub-and-spoke networking, branch connectivity, global routing, and Azure landing zone connectivity.
  • AWS Transit Gateway Documentation
    What it is: AWS documentation for Transit Gateway, a regional network transit hub for connecting VPCs and on-premises networks.
    Why it is useful: Essential for AWS enterprise network design, segmentation, routing domains, and large-scale VPC connectivity.
  • Google Cloud Network Connectivity Center
    What it is: Google Cloud documentation for hub-and-spoke network connectivity management.
    Why it is useful: Useful for designing Google Cloud hybrid and multicloud connectivity models.
  • Amazon VPC Documentation
    What it is: AWS documentation for Virtual Private Cloud networking.
    Why it is useful: Useful for subnet design, route tables, gateways, endpoints, security groups, NACLs, and cloud network isolation.
  • Azure Networking Documentation
    What it is: Microsoft’s central documentation hub for Azure networking services.
    Why it is useful: Useful for virtual networks, DNS, firewalls, load balancing, private endpoints, gateways, and routing.
  • Google Cloud VPC Documentation
    What it is: Google Cloud documentation for Virtual Private Cloud networking.
    Why it is useful: Useful for understanding Google Cloud’s global VPC model, subnets, routing, firewall rules, peering, and hybrid connectivity.
  • VMware NSX Documentation
    What it is: VMware NSX documentation for software-defined networking and security.
    Why it is useful: Critical for architects working with overlay networking, microsegmentation, distributed firewalls, edge routing, and private cloud network modernization.

Kubernetes, Containers, and Platform Engineering

These links are useful when designing container platforms, internal developer platforms, GPU-enabled Kubernetes, and cloud-native operations.

  • Kubernetes Documentation
    What it is: The official documentation for Kubernetes concepts, workloads, services, storage, networking, scheduling, and cluster administration.
    Why it is useful: Essential for architects designing container platforms, platform engineering services, AI clusters, and cloud-native application environments.
  • Docker Documentation
    What it is: Official documentation for Docker Engine, Docker Desktop, Compose, images, containers, registries, and container workflows.
    Why it is useful: Useful for understanding container packaging, local development workflows, image management, and container lifecycle basics.
  • Helm Documentation
    What it is: Official documentation for Helm, the Kubernetes package manager.
    Why it is useful: Useful for packaging, deploying, templating, and managing Kubernetes applications at scale.
  • Red Hat OpenShift Documentation
    What it is: Official documentation for Red Hat OpenShift Container Platform.
    Why it is useful: Helpful for enterprise Kubernetes design, cluster operations, developer platforms, GitOps, container security, and regulated workload hosting.
  • NVIDIA AI Enterprise Documentation
    What it is: NVIDIA documentation for enterprise AI software, AI frameworks, NIM microservices, GPU operators, drivers, and AI platform components.
    Why it is useful: Useful for architects designing GPU-backed AI platforms across data center, cloud, and edge environments.
  • NVIDIA GPU Operator Documentation
    What it is: NVIDIA documentation for automating GPU software lifecycle management in Kubernetes.
    Why it is useful: Important for Kubernetes-based AI platforms that need GPU drivers, device plugins, monitoring, and GPU scheduling.
  • CNCF Landscape
    What it is: A searchable landscape of cloud-native projects, products, and vendors.
    Why it is useful: Useful for comparing tools across Kubernetes, service mesh, observability, GitOps, security, networking, and platform engineering.
  • Backstage Documentation
    What it is: Documentation for Backstage, an open platform for building developer portals.
    Why it is useful: Useful for internal developer platforms, service catalogs, golden paths, software templates, and engineering self-service.

Infrastructure as Code and Automation

These links are useful for automating infrastructure, building repeatable deployments, reducing configuration drift, and standardizing cloud operations.

  • Terraform Documentation
    What it is: Official Terraform documentation for infrastructure as code, providers, state, modules, workspaces, and configuration language.
    Why it is useful: Essential for multicloud provisioning, repeatable infrastructure deployment, drift management, and modular architecture patterns.
  • Terraform Registry
    What it is: Terraform’s registry for providers and reusable modules.
    Why it is useful: Useful for finding official providers for AWS, Azure, Google Cloud, VMware, Kubernetes, and many infrastructure platforms.
  • OpenTofu Documentation
    What it is: Documentation for OpenTofu, an open-source infrastructure as code tool compatible with Terraform-style workflows.
    Why it is useful: Useful for teams evaluating open-source infrastructure as code options and Terraform-compatible workflows.
  • Pulumi Documentation
    What it is: Documentation for Pulumi, an infrastructure as code platform that uses programming languages like Python, TypeScript, Go, .NET, Java, and YAML.
    Why it is useful: Useful for architects who want infrastructure automation with general-purpose programming language patterns.
  • Ansible Documentation
    What it is: Community documentation for Ansible automation, playbooks, inventories, modules, and collections.
    Why it is useful: Useful for configuration management, server automation, network automation, and repeatable operational runbooks.
  • Red Hat Ansible Automation Platform Documentation
    What it is: Red Hat’s enterprise documentation for Ansible Automation Platform.
    Why it is useful: Useful for enterprise automation, workflow orchestration, role-based access, job templates, automation mesh, and governed automation programs.
  • Packer Documentation
    What it is: Documentation for Packer, a tool for creating machine images across platforms.
    Why it is useful: Useful for golden image pipelines, VM templates, cloud images, and standardized operating system builds.
  • Crossplane Documentation
    What it is: Documentation for Crossplane, a Kubernetes-based control plane framework for managing infrastructure and services.
    Why it is useful: Useful for platform engineering teams building self-service infrastructure APIs on top of Kubernetes.

DevOps and Delivery

These links are useful for CI/CD, GitOps, source control workflows, release engineering, and software delivery pipelines.

  • GitHub Actions Documentation
    What it is: GitHub’s documentation for automating workflows, CI/CD, testing, builds, and deployment pipelines.
    Why it is useful: Useful for automating infrastructure code, application delivery, compliance checks, documentation builds, and AI-assisted development workflows.
  • GitHub Docs
    What it is: GitHub’s documentation portal for repositories, pull requests, Actions, Codespaces, security, Copilot, and administration.
    Why it is useful: Useful for source control governance, developer productivity, secure coding workflows, and repository management.
  • GitLab CI/CD Documentation
    What it is: GitLab documentation for CI/CD pipelines, runners, jobs, stages, artifacts, and deployment workflows.
    Why it is useful: Useful for teams standardizing build, test, security scan, and deployment pipelines in GitLab.
  • Azure Pipelines Documentation
    What it is: Microsoft documentation for Azure Pipelines in Azure DevOps.
    Why it is useful: Useful for enterprise CI/CD, release pipelines, environment approvals, infrastructure deployments, and Microsoft ecosystem automation.
  • Argo CD Documentation
    What it is: Documentation for Argo CD, a GitOps continuous delivery tool for Kubernetes.
    Why it is useful: Useful for Kubernetes application deployment, declarative configuration, drift detection, and GitOps platform patterns.

Security, Governance, and Risk

These links are useful for AI governance, cloud security, LLM risk management, adversarial AI threat modeling, security controls, and architecture guardrails.

  • NIST AI Risk Management Framework
    What it is: NIST’s framework for managing risks associated with artificial intelligence.
    Why it is useful: Useful for AI governance, risk reviews, control mapping, executive communication, and responsible AI program design.
  • NIST AI RMF Resource Center
    What it is: NIST’s resource center for AI Risk Management Framework guidance, profiles, and supporting material.
    Why it is useful: Helpful when translating AI governance principles into practical review steps and control language.
  • OWASP Top 10 for LLM Applications
    What it is: OWASP’s security risk list for large language model applications.
    Why it is useful: Critical for identifying common LLM security risks such as prompt injection, data leakage, insecure output handling, excessive agency, and supply chain exposure.
  • OWASP Generative AI Security Project
    What it is: OWASP’s broader generative AI security project site.
    Why it is useful: Useful for AI application threat modeling, secure design checklists, and practical security education.
  • MITRE ATLAS
    What it is: MITRE’s knowledge base of adversary tactics and techniques against AI systems.
    Why it is useful: Useful for AI threat modeling, red-team planning, security architecture, and mapping attack techniques to defensive controls.
  • Cloud Security Alliance AI Controls Matrix
    What it is: CSA’s control framework for secure and responsible cloud AI systems.
    Why it is useful: Useful for mapping AI platform risks to controls, governance requirements, compliance conversations, and enterprise security programs.
  • Cloud Security Alliance Cloud Controls Matrix
    What it is: CSA’s cloud security control framework.
    Why it is useful: Useful for cloud security assessments, vendor reviews, control mapping, and governance design across multiple cloud providers.
  • CIS Benchmarks
    What it is: Security configuration benchmarks for operating systems, cloud services, Kubernetes, databases, and infrastructure platforms.
    Why it is useful: Useful for hardening baselines, compliance evidence, configuration reviews, and security automation.

AI Platforms and Model Providers

These links are useful for building enterprise AI applications, generative AI platforms, model services, agent systems, inference patterns, and AI architecture designs.

  • Microsoft Foundry Documentation
    What it is: Microsoft’s documentation for building, managing, evaluating, and governing AI applications and agents on Azure.
    Why it is useful: Useful for enterprise AI architects designing Azure AI platforms, model catalogs, agent systems, evaluation workflows, and governance.
  • Microsoft Foundry Portal
    What it is: Microsoft’s web portal for building and managing AI projects, models, agents, and related resources.
    Why it is useful: Useful for hands-on prototyping, model selection, prompt testing, agent workflows, and AI project management.
  • Amazon Bedrock Documentation
    What it is: AWS documentation for Amazon Bedrock, including models, inference, agents, knowledge bases, guardrails, and evaluation features.
    Why it is useful: Useful for designing AWS generative AI applications, model access patterns, retrieval workflows, and governed AI services.
  • Amazon Bedrock Product Page
    What it is: AWS’s product overview for Amazon Bedrock.
    Why it is useful: Useful for understanding service positioning, supported capabilities, business messaging, and high-level architecture fit.
  • Google Gemini Enterprise Agent Platform Documentation
    What it is: Google Cloud’s documentation hub for enterprise AI, agents, models, tools, and related APIs.
    Why it is useful: Useful for architects designing Google Cloud AI applications, agent systems, model workflows, and enterprise AI integrations.
  • OpenAI API Documentation
    What it is: OpenAI’s developer documentation for building with models, tools, agents, and APIs.
    Why it is useful: Useful for understanding OpenAI application patterns, API usage, tool calling, structured outputs, and agent design.
  • OpenAI API Reference
    What it is: OpenAI’s API reference for endpoints, request parameters, response objects, and developer implementation details.
    Why it is useful: Useful when building production integrations, validating schemas, debugging API behavior, and creating reusable AI services.
  • Hugging Face Documentation
    What it is: Documentation for Hugging Face Hub, models, datasets, transformers, inference, and related AI libraries.
    Why it is useful: Useful for model discovery, open model evaluation, dataset exploration, prototyping, and AI engineering workflows.
  • NVIDIA AI Enterprise Documentation
    What it is: NVIDIA’s enterprise AI software documentation.
    Why it is useful: Useful for GPU-enabled AI architecture, enterprise inference, model serving, Kubernetes AI platforms, and accelerated computing designs.

AI Agents, Orchestration, and Protocols

These links are useful for designing agentic AI systems, tool-using assistants, multi-agent workflows, AI integrations, and emerging interoperability patterns.

  • Model Context Protocol Documentation
    What it is: Documentation for MCP, an open standard for connecting AI applications to tools, data sources, and external systems.
    Why it is useful: Useful for architects designing reusable AI tool integrations, enterprise context access, agent connectors, and governed AI workflows.
  • Agent2Agent Protocol
    What it is: An open protocol for communication between AI agents.
    Why it is useful: Useful for understanding emerging multi-agent interoperability patterns and cross-system agent coordination.
  • Microsoft Agent Framework Documentation
    What it is: Microsoft documentation for building agentic AI solutions with its Agent Framework.
    Why it is useful: Useful for designing enterprise-grade agents in Microsoft environments using .NET, Python, orchestration, workflows, and tool integrations.
  • Microsoft Foundry Agent Service
    What it is: Microsoft’s managed agent service documentation within Foundry.
    Why it is useful: Useful for building, deploying, scaling, and governing agents on Azure.
  • Amazon Bedrock AgentCore Documentation
    What it is: AWS documentation for building, deploying, and operating AI agents with Bedrock AgentCore.
    Why it is useful: Useful for AWS-based agent architectures, agent runtime design, tool integration, memory, identity, and enterprise operations.
  • Google Agent Development Kit Documentation
    What it is: Google’s documentation for its Agent Development Kit.
    Why it is useful: Useful for building, testing, and deploying agent systems in Google Cloud’s AI ecosystem.
  • OpenAI Agents SDK Documentation
    What it is: Documentation for OpenAI’s Python Agents SDK.
    Why it is useful: Useful for building tool-using agents, multi-step workflows, handoffs, guardrails, and traceable AI applications.
  • OpenAI Agents Guide
    What it is: OpenAI’s guide for designing and implementing agents with the OpenAI platform.
    Why it is useful: Useful for understanding core agent architecture patterns, tools, tracing, and implementation choices.
  • Semantic Kernel Documentation
    What it is: Microsoft documentation for Semantic Kernel, an AI orchestration SDK for agents and AI applications.
    Why it is useful: Useful for architects building AI orchestration patterns, planners, tools, memory, and enterprise integrations.

RAG, Vector Databases, and Knowledge Systems

These links are useful for retrieval-augmented generation, knowledge ingestion, embeddings, semantic search, vector databases, document parsing, and grounding AI systems in enterprise data.

  • LlamaIndex Documentation
    What it is: Documentation for LlamaIndex, a framework for building knowledge-augmented AI applications.
    Why it is useful: Useful for RAG pipelines, document ingestion, indexing, retrieval, agents, and connecting LLMs to enterprise knowledge.
  • LangChain Documentation
    What it is: Documentation for LangChain tools, integrations, agents, workflows, and LLM application development.
    Why it is useful: Useful for prototyping and building LLM applications that connect models, prompts, tools, retrievers, and external systems.
  • Pinecone Documentation
    What it is: Documentation for Pinecone, a managed vector database for semantic search and AI retrieval.
    Why it is useful: Useful for RAG architectures, similarity search, retrieval performance, vector indexing, and production knowledge systems.
  • Qdrant Documentation
    What it is: Documentation for Qdrant, a vector search engine and vector database.
    Why it is useful: Useful for semantic search, open-source vector storage, filtering, embeddings, and AI retrieval architectures.
  • pgvector
    What it is: An open-source PostgreSQL extension for vector similarity search.
    Why it is useful: Useful when teams want vector search inside PostgreSQL instead of introducing a separate vector database.
  • Weaviate Documentation
    What it is: Documentation for Weaviate, an open-source vector database.
    Why it is useful: Useful for semantic search, hybrid search, embeddings, RAG applications, and knowledge graph style AI retrieval.
  • Chroma Documentation
    What it is: Documentation for Chroma, an open-source AI data and vector database project.
    Why it is useful: Useful for local prototyping, RAG development, embeddings, and lightweight vector search workflows.
  • LlamaParse Documentation
    What it is: Documentation for LlamaParse, a document parsing service from LlamaIndex.
    Why it is useful: Useful for converting complex PDFs, documents, tables, and enterprise content into AI-ready data.

MLOps, LLMOps, Evaluation, and Testing

These links are useful for tracking experiments, evaluating models, testing prompts, monitoring AI behavior, and building repeatable AI delivery pipelines.

  • MLflow Documentation
    What it is: Documentation for MLflow, including experiment tracking, model registry, evaluation, tracing, prompt management, and AI governance workflows.
    Why it is useful: Useful for MLOps and LLMOps teams that need repeatable model development, testing, evaluation, governance, and lifecycle management.
  • Promptfoo Documentation
    What it is: Documentation for Promptfoo, an open-source tool for LLM testing, evaluation, red teaming, and benchmarking.
    Why it is useful: Useful for validating prompts, models, RAG systems, agent behavior, safety checks, and regression testing before production.
  • Hugging Face Documentation
    What it is: Documentation for Hugging Face models, datasets, libraries, inference, and AI tooling.
    Why it is useful: Useful for model comparison, open-source model testing, dataset discovery, fine-tuning research, and AI prototyping.

API Design and Integration

These links are useful for designing APIs, documenting services, building integration patterns, and connecting cloud and AI systems together.

  • OpenAPI Specification
    What it is: The OpenAPI Specification for describing HTTP APIs in a standard format.
    Why it is useful: Useful for documenting APIs, generating clients, validating contracts, supporting integration patterns, and enabling AI tools to understand API behavior.
  • OpenAPI Learn
    What it is: Educational documentation for learning OpenAPI concepts and best practices.
    Why it is useful: Useful for architects and developers who need to standardize API design and documentation practices.
  • AsyncAPI Documentation
    What it is: Documentation for AsyncAPI, a specification for event-driven APIs and messaging systems.
    Why it is useful: Useful for event-driven architectures, streaming platforms, pub-sub systems, message brokers, and asynchronous integrations.
  • Swagger Documentation
    What it is: Documentation for Swagger tools used with OpenAPI.
    Why it is useful: Useful for API design, interactive documentation, API testing, and developer enablement.
  • Postman Documentation
    What it is: Documentation for Postman API design, testing, collections, environments, automation, and collaboration.
    Why it is useful: Useful for API validation, integration testing, documentation, mock APIs, and team collaboration.
  • OpenAPI Tools
    What it is: A directory of tools that support OpenAPI workflows.
    Why it is useful: Useful for finding generators, validators, documentation tools, testing tools, and API lifecycle utilities.

Observability and Incident Response

These links are useful for telemetry, metrics, logs, traces, dashboards, alerting, SRE practices, and production operations.

  • OpenTelemetry Documentation
    What it is: Documentation for OpenTelemetry, a vendor-neutral observability framework for traces, metrics, and logs.
    Why it is useful: Useful for designing portable telemetry across cloud, Kubernetes, microservices, AI applications, and vendor observability tools.
  • Prometheus Documentation
    What it is: Documentation for Prometheus, an open-source monitoring and alerting toolkit.
    Why it is useful: Useful for Kubernetes metrics, infrastructure monitoring, service health, alerting rules, and platform observability.
  • Grafana Documentation
    What it is: Documentation for Grafana dashboards, observability, incident response, alerting, and related tools.
    Why it is useful: Useful for visualizing metrics, logs, traces, SLOs, platform health, and operational trends.
  • Elastic Observability Documentation
    What it is: Elastic documentation for observability across applications, infrastructure, logs, metrics, traces, and synthetic monitoring.
    Why it is useful: Useful for centralized logging, search, operational analytics, infrastructure monitoring, and troubleshooting.
  • Azure Monitor Documentation
    What it is: Microsoft documentation for Azure Monitor, logs, metrics, alerts, workbooks, and application insights.
    Why it is useful: Useful for monitoring Azure workloads, hybrid resources, applications, infrastructure, and platform health.
  • Amazon CloudWatch Documentation
    What it is: AWS documentation for CloudWatch metrics, logs, alarms, dashboards, events, and observability features.
    Why it is useful: Useful for AWS workload monitoring, operational alarms, log analysis, and infrastructure visibility.
  • Google Cloud Observability Documentation
    What it is: Google Cloud documentation for monitoring, logging, tracing, profiling, and error reporting.
    Why it is useful: Useful for operating Google Cloud workloads, troubleshooting production issues, and building cloud observability practices.

FinOps, Pricing, and Cost Management

These links are useful for estimating costs, managing cloud spend, building FinOps practices, and explaining financial impact to business stakeholders.

  • FinOps Framework
    What it is: The FinOps Foundation’s operating model for cloud financial management.
    Why it is useful: Useful for creating a common language between engineering, finance, procurement, and business teams.
  • Azure Pricing Calculator
    What it is: Microsoft’s calculator for estimating Azure product and service costs.
    Why it is useful: Useful for project estimates, architecture tradeoffs, migration planning, and executive cost conversations.
  • AWS Pricing Calculator
    What it is: AWS’s calculator for estimating AWS service costs.
    Why it is useful: Useful for workload planning, cost comparisons, migration assessments, and AWS architecture pricing.
  • Google Cloud Pricing Calculator
    What it is: Google Cloud’s pricing calculator for estimating product and service costs.
    Why it is useful: Useful for Google Cloud workload estimates, design comparisons, and cost planning.
  • Broadcom VMware License Calculator
    What it is: Broadcom’s knowledge base article for the VMware Cloud Foundation, VMware vSphere Foundation, and vSAN license calculator.
    Why it is useful: Useful for private cloud planning, VMware licensing discussions, capacity analysis, and cost modeling.
  • Azure Cost Management and Billing Documentation
    What it is: Microsoft documentation for Azure cost management, billing, budgets, exports, and optimization.
    Why it is useful: Useful for cost governance, budget alerts, showback, chargeback, and cloud spend optimization.
  • AWS Cost Management Documentation
    What it is: AWS documentation for cost management, billing, budgets, cost explorer, and optimization tools.
    Why it is useful: Useful for AWS financial governance, spend visibility, usage analysis, and optimization programs.
  • Google Cloud Billing Documentation
    What it is: Google Cloud documentation for billing, budgets, cost controls, exports, and cost management.
    Why it is useful: Useful for managing Google Cloud spend, analyzing usage, and integrating billing data into FinOps workflows.

Learning, Labs, and Hands-On Practice

These links are useful for continued learning, certifications, hands-on labs, cloud fundamentals, AI skills, and platform engineering growth.

  • Microsoft Learn Azure Training
    What it is: Microsoft’s training portal for Azure learning paths, modules, labs, and certification preparation.
    Why it is useful: Useful for building Azure skills across architecture, administration, AI, security, data, and development.
  • AWS Skill Builder
    What it is: AWS’s cloud learning platform with courses, labs, exam prep, and role-based training.
    Why it is useful: Useful for developing AWS architecture, AI, security, operations, networking, and migration skills.
  • AWS Digital Training
    What it is: AWS’s digital training page for online cloud courses and learning resources.
    Why it is useful: Useful for structured AWS learning, certification preparation, and role-specific development.
  • Google Skills
    What it is: Google’s skills platform for cloud, data, AI, and technical learning.
    Why it is useful: Useful for building Google Cloud and AI skills through guided learning and practical exercises.
  • VMware Hands-on Labs
    What it is: VMware’s hands-on lab platform for exploring VMware products and solutions.
    Why it is useful: Useful for practicing VMware, VCF, NSX, vSphere, Aria, and cloud platform scenarios without building a full lab environment.
  • VMware Hands-on Labs Catalog
    What it is: The VMware Hands-on Labs catalog.
    Why it is useful: Useful for quickly finding labs by product, topic, or learning objective.
  • Kubernetes Tutorials
    What it is: Official Kubernetes tutorials for hands-on learning.
    Why it is useful: Useful for strengthening practical Kubernetes skills around workloads, services, storage, configuration, and cluster operations.

How I Recommend Using This Page

  • For architecture design: Start with the architecture centers and Well-Architected frameworks.
  • For implementation: Use the official product documentation, CLI references, and infrastructure as code links.
  • For AI platform work: Start with Microsoft Foundry, Amazon Bedrock, Google Gemini Enterprise Agent Platform, OpenAI, Hugging Face, and NVIDIA AI Enterprise.
  • For AI security: Use NIST AI RMF, OWASP LLM Top 10, MITRE ATLAS, CSA AI Controls Matrix, and CIS Benchmarks.
  • For operations: Use OpenTelemetry, Prometheus, Grafana, Elastic Observability, Azure Monitor, Amazon CloudWatch, and Google Cloud Observability.
  • For business alignment: Use FinOps, pricing calculators, cloud adoption frameworks, and cost management documentation.

Disclaimer: External documentation changes frequently. Always validate product versions, service availability, limits, licensing, and regional support directly with the vendor before making production architecture decisions.