Implementing Observability in K8s | DataOps Day
No production system is complete without a way to monitor it. In software we define observability as the ability to understand how our system is performing. This talk dives into capabilities and tools that are recommended for implementing observability when running K8s in production as the main platform today for deploying and maintaining containers with cloud-native solutions. We start by introducing the concept of observability in the context of distributed systems such as K8s and the difference with monitoring. We continue by reviewing the observability stack in K8s and the main functionalities. For example, we could review the K8s control plane for managing the cluster state. Finally, we will review the tools K8s provides for monitoring and logging, and get metrics from applications and infrastructure.
Between the points to be discussed we will highlight:
-Introducing the concept of observability
-Observability stack in K8s
-Tools and apps for implementing Kubernetes observability
-Integrating Prometheus with OpenMetrics
Transcript
Uh, good morning everyone. My name is Homan Orga. Uh, this talk, uh, I'm going to talk about implementing ulcerative for Kubernetes.
Uh, this is the agenda for representation. We start with by intrusion the course, the, the concept of ity in the context of a distributed systems, uh, such as Kubernetes and the defense with monitoring. We continue re reviewing the, the Solidity stack in Kubernetes and the main functionalities.
And finally, we'll review the tools, uh, Kubernetes provides for monitoring and logging and get metrics from, from application and infras and infras infrastructure. Well, in today's, uh, area of microservices, uh, and, and application. So what a ture is more complex and as that tors meet demand, uh, for high, uh, ability, right, reliability, security, and auto scaling during big time.
The scale of, eh, this operation requires robust management and the three pillars of observability, like logs, metrics, and traces to the book and maintain systems and application. In this way, coordinators observability is no, uh, a top, uh, for, for, for, for DevOps teams, observability. What is observ?
Ally is property of the IT ecosystem that provides the infrastructure, the comprehensive and and understanding of assisting infrastructure, uh, through metrics, logging and tracing. Basically, monitoring allows, uh, to extract metrics from your applications and resources that can later be visualized and analyzed to present the, the current state of your or your resources. Uh, once the metrics are, are restricted, then uh, they can be used to set up all rules, facilitating, profiling and devo tax.
Uh, and the second, uh, one, uh, we have a login that allows the developers to the book they containers in case of failure, who is, for example, Docker provide an active way to approaching container locks, but it is very limited in, in, in its functionality. As a result, a centralized lock, a plane is AMAs in any container oriented environment. And finally, RAC, that allow us to de book services running, uh, on a network and follow a request trade.
Until there, the source, the source of a problem can be determined. Well, a metric, what is a metric, basically metric is the fundamental type of, uh, signal that can be emit by a service or the infrastructure is running on, is basically a metric, is the combination of some of some identifier that indicates what the metric represents, and a series of data points, each of which contains two elements, the time stands in which the data point was generated, and an made value representing the state of the thing you are measuring. And, uh, at that time, regarding enticing, we can define traceability as the ability to know exactly where each operation carry out in the system goes, goes, uh, go through from each origin to its end.
Uh, trace basically is a representation of consecutive events, which reflect an end to end request flow in, uh, in systems. The following D diagram shows this, the trace representation of our request flow. Uh, trace basically detected a cyclical graph of span where the between spans are called references.
In Kubernetes, we have, uh, uh, uh, the following key metrics to monitor, uh, Kubernetes clatters, for example, monitoring Kubernetes involve inspecting all the core components like clusters, no ports, deployments and services, and cluster monitoring. Health administrator understand that, uh, water resources are used by the cluster in the cluster for capacity planning and the number of applications running on each, uh, on each another. This, eh, in this is like you, we can see the, the main metrics we can use for monitoring, uh, a Kubernetes cluster, for example, if we need monitoring, uh, PO for support, PO level, these are the main metrics.
For example, we have Kubernetes metrics, uh, container metrics and application metrics. E Each one of these metrics apply can apply, uh, for monitoring the, the resources of your application, like say, C P U, memory network usage, uh, the number of, uh, requests are your application is, is doing, uh, and so on the system, resources and so on. Basically, when we, when we want to implement monitoring and observating Kubernetes, we need to mo to monitoring these components.
The main components like, uh, the Kubernetes, uh, the BER server, the cover controller, control manager, the cellular, the CT c D database, and so on. For a starting, we call it start, uh, with passive monitoring with in your, in our CNET cluster with ES dashboard, which you can use, for example, this dashboard to get an overview of applications running on your cluster. Uh, as well as for creating, uh, or modifying the, uh, Kubernetes resources such as deployment jobs and diamond sets.
Kubernetes tools are so solutions that help trial the performance, health and resource of Kubernetes cluster, Northern and containers. These tools that are, are we going to review, are the main tools that provide insights, uh, into various components of a Kubernetes environment in na and enable administrators and developers to maintain and optimize their applications. And next, we're going to review, uh, how to implement our strategy using open source to, uh, technologies.
For example, Kuku Watch is an open source Kubernetes monitoring tool that sends notifications about changes in an Kubernetes cluster to many communication channels. Um, basically it had the capacity to monitoring Kubernetes resources such as deployment services and bots. Another interesting tool is Loop.
Loop is a solution that monitors stores and visualize events and change in Kubernetes resources Over time. It is just sign it, uh, to provide a timeline of data, eh, made to existing resources and resources that no longer exist in in the cluster. The visual dashboard swallows for ac spectrum of event metrics for the booking, uh, or error harling proposed.
We continue with Jagger Jagger, open source distributed sistance. This tool is descending to monitor and trouble salt distributed microservices, mostly focusing on distributed content, pro uh, con context preparation, distributed transaction monitoring, lu cow analysis, service dependency analysis, and performance latency optimization. This tool, uh, provides a client liberties, um, for most programming language like Java, go and Python.
And the idea is that your applications needs to be instrumented before Inca. Inca can set information. This tool is very useful for Red Roadhouse analysis by the soliciting traces and identifying both bottleneck or errors in, in the system, uh, at application or infrastructure level.
And the developers can perform a root cow analysis to optimize the, their applications, uh, and improve overall performance. For example, we could quickly see the requests that are being made and, uh, and the duration to the tech bottlenecks or requests who, uh, duration can be atypical or anomalous in the application can be deployed in Kubernetes jaar operator, which helps to deploy and manage the Jaar instances. The simplest way to create a Jaar instance is by creating a configuration file, uh, deploying the all in one image, convening the jaar, uh, the Jaar agent, the Jaar collector, the Jaguar query, and the Jaar, uh, using an interface in a simple, in a simple pot, using in memory storage by the default.
Another interesting, uh, tool is flu nd that allow us to working with Kubernetes, with flu and d uh, open source data collector for unified logging layers. It was with Kubernetes running as demonstrate, and this combination ensures that all no drew one copy of a pot. This example shows, uh, the configuration, the Demonstr configuration manifests with deploys a flu D agent using the container image.
Fluent de es demonstrate on every ES node, uh, in the, in the cluster flu D also had the capacity to supporting multiple, uh, data output plugins, uh, for export logs to 30 party applications like Elastic Search, for example. You can integrate it, uh, with a cluster, uh, uh, that, that provides elastic. And if your application, uh, running Kubernetes is a streaming logs to a standard, uh, output, uh, the flu, the flu, the flu ND engine again, uh, will ensure everything is develop correctly to a centralized logging solution.
Another, uh, tool for review and at this, at this, uh, point is material that is a ative, uh, time series data, uh, that the store with store with ING Rich query language for metrics, collecting data with pro materials open as many possibilities for increasing the disability of genome infrastructure and the containers running in the, in the Kubernetes clo, uh, cluster. Kubernetes Pro, uh, provides, uh, a multidimensional data models that use at time series data model with metric names and, uh, base current levels enable flexible and powerful query. Also, also provide a, a query language that allows voices to aggregate, filter and manipulate collective metrics for analyzing and alerting purpose, and also provides another interesting, uh, filters like data collection, uh, provides, uh, uh, an additional storage and visualization at this point.
Uh, highlight that, uh, usually pro itegrated with Grafana for visualization tax. The following diagram shows the high level of, of promoters and its external components, the pro of services discovery, as well as pulling metrics from, from the sport test to be storing the material database, database and ladder use for visualization or, or, uh, or alerting. Alert manager is, is responsible for setting up all rules, analyzing the data in the database and sending alert messages, uh, to multiple receivers in a, a certain rule is trigger.
And finally, the other component that use, uh, that is used for proven use is exported, that are standalone and independent processes that can be executed on our target resource to generate and export metrics via metrics a p i in a production, um, pro, uh, Kubernetes environment. The following metrics, uh, should be suitable in, in your dashboard and can be exported using, for example, not exporter and visor. Uh, these are the main metrics that we can, that, that we can use for, for, for measure.
The, the resource, uh, for check the nu matter of filing post and error with a specific, uh, name space or the ES resource capacity. That is the total number of nodes C P U course and Memorial available in your, in your cluster. For example, C has become one of the leading solutions for monitoring containers because of each user friendliness, flexibility, and ability to meet atmos and more any monitoring requirement with Visor, uh, you can, you could collect aggregate process and export container, uh, by as such as C P U and memory usage, file systems and network statistics.
And finally, the, uh, another intestine tool that we have for, for, for, for, for getting in in Kubernetes is that is a fully distributed and security observability platform for ative, uh, work logs. It is boil on top of CION and E B P F to enable DB visibility into the communication and behavior of service of services, as well as, uh, the working infras in a completely transparent way, basically, uh, can ask her questions such as what services are communicating with each other? What H D P calls are being made is an enable communicating failing?
Is the communication broken at, at a on a specific layer? Or, for example, which services have connection block due to network policy? These are the question that we can answer using this.
The, this tool, the basis of Hubbell is, uh, the use of technologies such as cion as E V P F. Cilium is an open source ative solution for providing, uh, securing and observing network connectivity between, uh, word logs and the E V P F. Uh, the technology allows to give the facility of the systems and applications with a granularity and improved levels of efficiency.
Habell enables for automatic discovery of the service dependency graph for Kubernetes cluster, allowing user-friendly visual decision and filtering of those data flows as a service map. The metrics and monitoring functionality provides an overview of the state of applications and allow to recognize patterns indicating failure and other scenarios that require, uh, action also allows Pro also provides this tool also provide visibility for into flow information on the network and and application protocol level. This enables visibility into individual T C P connections, uh, queries and HT p requests.
Finally, by, I'm going to comment how we can integrate, uh, Prometheus as a tool for, uh, getting meters with open telemetry. Open Telemetry is a solution BA based on signal on signals, emit a might em meter by applications and provides a number of components to implement observability. And it's a tool very useful for, for implementing observability in, in, in many, uh, many environments including, uh, including Kubernetes.
They seen that an application can emit basically are, uh, the traces, metrics and logs in the open telemetry infrastructure architecture. Uh, we can see how there is a series of microservices responsible of for managing the observability information and sending in to, to the collector component offered a offered by a open telemetry. Open telemetry is based of the use and the, on the use of a collector that the co descending of the traces, uh, to the different of ity backends from the application.
Tele eliminating problems associated with errors, and the co the configuration of the collector is done with a jam configuration file, and it's based on the, the, these three many elements. You have receivers that are the, that are source of OSS information processors that processed information received be before, before it export to the different backs and exported, exported that they are in change, uh, responsible, then they are responsible of exporting the information to different bucket such as ger. For example.
In this diagram, we can see microservice and microservice architecture where for each service, the Open telemetry client, a company is added to collect the, the shareability information for each micro service and send in to the collector, which in, in, in this case is jaar. Jaar have the capacity to collect of the, of the, of the metrics, uh, for, for each service. And in a, in a, in a, in a uniform repository, this could be the collector configuration for image service that allows collect, uh, the shareability, the, the observability information for each one.
Basically with the, this, the, this configuration, the, we define the image, the command to secure the volume, the volumes, ports, sports, uh, port, and what is the image we can, what, what is the service, uh, uh, that we are, we are going to use for exporting this, uh, that, that open telemetry is going, is, is going to use for exporting this information for the collector. We also, we need the following configuration for sending information to Jager. In this case, we can see, uh, not the pipelines inside the service, uh, that indicates how the flow of receivers, processors and exporters, uh, will be for the traces.
In this case, receivers re in the receivers, uh, that we are using, uh, the open telemetry standard port, uh, 4 4 3 1 7, the, and in the professors we can see that, uh, we are using the batch, uh, professors. In this case, we indicate that it will be in batch model. That is, it will send batches of information to, to Jager.
And in the supporter sector, we export the information, uh, to, to the Jaar, to the Jager SE service in the following repository. We consider the different professors we call use, we call use with open telemetry. Professors are used at various stages for a pipeline.
Generate a Professor prepro data before each export or health, ensure that data makes it through a pipeline successful. In the same way we have used Jaar as a supporter, we could use other backend services like Prometheus or Grafana as backend metrics. This could be the Prometheus configuration in the service.
Just sector. In this configuration, we are refining according to the defining interval, uh, the metric endpoints of each service to retrieve the, the information. And finally, uh, in the, in the exporter sector, uh, we call define the, the, in the past, the post exporter to pro use, uh, to make it worse is necessary, uh, to activate this feature in the Protive server.
With the Optium web enabled remote right receiver, you can experiment with open telemetry, uh, with the following repository they provide. In this, in the repository, you can find, uh, an open telemetry configuration, uh, help charts to deploy the application to an existing Bernet cluster. In this view, we can, uh, this, in this match, you can see the ion of this, uh, open telemetry demo.
In this view, you can find the filter traces. For example, you, you can visualize the traces for a specific service. If you select, uh, one of the traces, you can see the different spans going through multiple services.
And if we take look at the, at the receive metrics, you should see a lot of different metrics. For example, you can filter using the course total text to see the metric created, uh, by the span, uh, metrics professor. And finally, for conclusions, uh, uh, well, at this pointing for concluding, uh, it is important to, to, to take account that, uh, lie of the lean on the native capabilities of Kubernetes for the collection and exploitation of metrics in order to know the state of hell of your post and in general of your cluster.
And, uh, all these, all the tools that I, I, I have reviewed, uh, you can use all these tool and the metrics that provide these tools to be, to be able to create alerts that pro proactively notify us of errors or even allow us to anticipate usage in your, in our application or infrastructure. And that's all. Thank you.
Thank you very much. Uh, you can contact me. You have a question, uh, for the presentation.
You can, uh, uh, contact me and you can do it via Twitter, for example, or LinkedIn. And also you can check my personal website where you can see, uh, other conference and other books. For example, I have published some books related with DevOps and DevSecOps in Dock and Kubernetes.
And, and you can find my, my personal page for, for checking this, uh, this information and for my parties. Uh, all uh, thank you for your attention. Uh,





