Sira Consulting

Victoria Metrics Architect

⭐ - Featured Role | Apply direct with Data Freelance Hub
This role is for a "Victoria Metrics Architect" with a contract length of "unknown" and a pay rate of "unknown." Key skills include hands-on experience with Victoria Metrics, observability platforms, Kubernetes, and security architecture. Telecommunications industry experience is preferred.
🌎 - Country
United States
💱 - Currency
$ USD
-
💰 - Day rate
Unknown
-
🗓️ - Date
July 21, 2026
🕒 - Duration
Unknown
-
🏝️ - Location
Unknown
-
📄 - Contract
Unknown
-
🔒 - Security
Unknown
-
📍 - Location detailed
United States
-
🧠 - Skills detailed
#Time Series #Prometheus #GIT #Base #GitLab #Storage #Stories #Grafana #Jira #Kubernetes #Scrum #Security #Deployment #Strategy #Leadership #HBase #Agile #"ETL (Extract #Transform #Load)" #Compliance #Observability #Project Management #Automation #Vault #Logging #Cloud
Role description
Victoria Metrics: Technical Skills Hands-on operation of Victoria Metrics in production — VM Cluster topology (VM Insert, VM Storage, VM Select), vm agent scrape and stream aggregation, VM Auth multi-tenant routing, VM Alert and VM Alert manager rule design, and Metrics QL query authoring. Ability to govern active time series cardinality, configure retention and down sampling, and design multi-cluster federation with cross-cluster deduplication and write-path failover. Experience Profile: Verifiable production VM Cluster deployments handling sustained, high-cardinality workloads — not evaluation or sandbox only. Evidence of diagnosing and remediating a cardinality problem in a live environment. Operational experience of VM Auth-based per-tenant data isolation beyond authentication alone. Familiarity with the Victoria Metrics Operator on Kubernetes and its upgrade lifecycle. Intermediate accepted where Expert-level Observability is demonstrated. Domain Level Technical Skills Experience Profile Observability Expert Expert-level design and operation of full-stack observability platforms: Prometheus scrape and federation, Grafana dashboard and datasource provisioning at scale, SLI/SLO definition and error-budget alerting, structured logging pipelines, and distributed tracing integration. Ability to define metric taxonomies and labelling standards that remain coherent across heterogeneous multi-vendor sources. Delivered observability platforms serving production engineering or operations teams in a large-scale environment. Designed SLO-based alerting frameworks that measurably reduced operational noise and accelerated incident response. Led metric standardisation across multiple teams or vendor sources. Evidence of instrumenting the observability platform itself — not only the workloads it serves. Architecture Leadership Expert Ability to serve as the single accountable design authority for a complex technical platform — producing High-Level Designs, Low-Level Designs, Architecture Decision Records, sequence diagrams, and interface contracts that engineering teams build from directly. Comfortable presenting architecture decisions and their trade-offs to mixed audiences from senior engineers to executive sponsors, adjusting depth without losing accuracy. Has held a Principal or Staff Architect role where design and build were explicitly separated. Maintained an ADR record throughout a programme delivery. Can name a design decision they were challenged on, the trade-offs documented, and how the decision held up through delivery. Evidence of architecture governance participation: design review boards, security assurance gates, and formal approval before engineering investment began. Security & Identity Expert Zero Trust architecture as a design default — every inter-component path authenticated, encrypted, and auditable. Deep integration with HashiCorp Vault including the Vault Secrets Operator synchronisation pattern, dynamic secrets, and PKI secrets engine for automated certificate issuance. Enterprise PKI lifecycle: certificate generation, renewal automation, CA distribution to distributed cluster workloads, and expiry alerting. OIDC and OAuth2 federation eliminating all static credential and local account patterns. Integrated Vault with Kubernetes workloads in production using the VSO sync pattern, not solely as a secrets store. Designed and operated automated TLS certificate lifecycle across a distributed platform including renewal without service interruption. Implemented OIDC federation for platform services against an enterprise identity provider. Evidence of security governance contribution: compliance reviews, threat model participation, or security assurance gate ownership. Kubernetes & OpenShift Intermediate Kubernetes administration and troubleshooting across StatefulSets, PersistentVolumeClaims, StorageClasses, Operators, RBAC, and NetworkPolicy. Red Hat OpenShift-specific competency: Security Context Constraints, OCP upgrade path management, OLM operator lifecycle, and the material differences from cloud-managed Kubernetes that affect stateful, high-throughput workloads. Multi-cluster topology design including hub and spoke architectures and cross-cluster connectivity. Operational experience on Red Hat OpenShift in an on-premises or bare-metal enterprise environment — not exclusively cloud-managed Kubernetes. Has configured SCCs for a stateful workload without cluster-admin workarounds. Managed an OCP cluster through at least one version upgrade including Operator compatibility validation. Experience with OpenShift ACM or equivalent for policy and workload deployment across a cluster fleet. GitOps & CI/CD Intermediate GitOps-native delivery model using ArgoCD or FluxCD as the reconciliation engine — all cluster state managed in Git, no manual production changes permitted. Kustomize overlay strategy for multi-environment and multi-tenant deployments: base resource definitions with environment-specific patches. GitLab CI/CD pipeline design including Kubernetes manifest validation gates, custom resource health checks, and environment promotion from lab through staging to production. Has operated a GitOps-exclusive delivery model in production where ArgoCD or Flux managed all cluster state and direct kubectl commands to production were prohibited. Designed a Kustomize overlay structure for a multi-cluster platform workload — not a single-cluster application. Built a GitLab CI pipeline with schema validation and promotion gates blocking invalid configurations before staging. Evidence of GitOps discipline applied to operator-managed custom resources, not only standard Deployments. Networking Intermediate Kubernetes networking spanning Ingress controllers, LoadBalancer service patterns, CoreDNS configuration, and NetworkPolicy for multi-tenant namespace isolation. MetalLB deployment on bare-metal OpenShift in BGP mode: IPAddressPool and BGPAdvertisement CRD configuration, eBGP peering with upstream ToR and spine switches, AS number design, prefix filtering, route-map policy, and BFD fast-failover. OVN-Kubernetes SDN design including EgressIP, EgressFirewall, and cross-cluster traffic segmentation. Deployed MetalLB in BGP mode on bare-metal Kubernetes or OpenShift and configured eBGP peering with physical switching infrastructure — not Layer 2 mode only. Designed NetworkPolicy rules for cross-namespace and cross-cluster traffic paths. Has operated in an IPv6 or dual-stack network environment. Evidence of DNS design for platform services: split-horizon resolution, external DNS automation, and service record lifecycle management. Telecommunications Intermediate Familiarity with the FCAPS operational framework — Fault, Configuration, Accounting, Performance, and Security — and how it shapes metric taxonomy, alerting rule design, and NOC-facing dashboard structure. Understanding of heterogeneous vendor telemetry integration: Prometheus exporter compatibility, OpenMetrics format validation, and label standardisation across multi-vendor network equipment from vendors such as Nokia and Ericsson. Has designed or operated an observability platform within a telecommunications, carrier-grade, or large-scale network infrastructure environment — 5G core, RAN, transport, or carrier edge. Experience normalising telemetry from heterogeneous vendor sources at the ingestion layer. Has integrated platform alerts into an enterprise NOC workflow including deduplication, suppression, and ticketing system integration. Prior engagement with a Tier-1 carrier, MVNO, or national-scale network operator strongly preferred. Programme Delivery Intermediate Fluent in agile delivery practices within a structured programme: sprint planning, backlog refinement, and epic and story decomposition from architecture into engineering tasks, maintaining architecture decisions ahead of sprint execution. Able to work effectively across Product Management, Project Management, and Scrum Masters — translating technical design choices into programme risk, timeline, and trade-off language that non-technical stakeholders can act on. Has operated as an architect embedded in a governed programme with sprint ceremonies, formal review gates, and stakeholder reporting obligations — not a solo or small-team context. Authored Jira epics and stories that engineering teams executed without requiring repeated design clarification. Has presented a technical design decision to a Product Manager or Project Manager and adapted framing based on their concerns. Evidence of architecture governance discipline under delivery pressure: ADRs completed, gates honoured, exceptions documented. Executive & Strategic Leadership Expert Developed and delivered enterprise IT strategy at Director level or above — translating business priorities into technology roadmaps, operating models, and investment cases for executive leadership. Experience leading large-scale digital transformation programmes spanning cloud adoption, platform modernisation, and operating model change. Partnered with C-Suite and SVP-level stakeholders to drive technology investment decisions, bridging business strategy and technology delivery across organisational boundaries. Has held a Director or equivalent role with enterprise-wide technology accountability — budget ownership above $50M, vendor governance, and SLA management across a global infrastructure estate. Built or led a Cloud Centre of Excellence delivering hybrid cloud governance and architectural standards. Led or governed a major managed services engagement (>$100M) including commercial model and operational improvement. Links technology investment to measurable business outcomes at programme scale.