Saturday, July 4, 2015

VMware Memory Management

There are four different methods by which ESX reclaims virtual machine memory. They are:
  • Ballooning
  • Transparent Page sharing
  • Hypervisor swapping
  • Memory compression
  •  

 Ballooning 


 it is a memory reclamation  technique used by a hypervisor to allow the physical host system to retrieve unused memory from certain guest virtual machines (VMs) and share it with others.

VMware memory ballooning, Microsoft Hyper-V dynamic memory, and the open source KVM balloon process are similar in concept.

Ex: if all the VMs on a host are allocated 8 GB of memory, some of the VMs only using half of the memory 4 GB. If one VM needs 12 GB of memory for an intensive process. Memory ballooning allows the host to borrow that unused memory and allocate it to the VMs with higher memory demand.  

Memory Ballooning with Real- Time Example

You are running a virtual Machine called ” VM1″ and You are starting a application called Microsoft Excel on that VM.

Windows will get the memory to run the application from the Guest physical memory.

Hypervisor sees the request and it get the memory from the Host physical memory .

After finish working with the application. Memory which is used is freed BUT since the hypervisor does not have access to Windows’ “free memory” list the memory will still be mapped in “host physical memory”

In case of an ESXi host running low on memory the hypervisor will ask the “balloon” driver installed inside the virtual machine (with VMware Tools) to “inflate”

In case of an ESXi host running low on memory the hypervisor will ask the “balloon” driver installed inside the virtual machine (with VMware Tools) to “inflate”

By default, Balloon driver (vmmemctl.sys) can reclaim upto a maximum of 65 % of guest physical memory. For example, You VM is allocated with 1000 MB of memory, It can be reclaimed upto 650 MB using this technique.
  
Analyzing Memory Ballooning Statistics:

You can verify the memory ballooning  stats from Esxtop ,Virtual Machine Resource Allocation tab and also using vCenter Performance Graphs.

esxtop -> Press m

You will see the “MEMCTL/MB” counter which shows us the overall ballooning activity (22110 MB). The “curr” and “target” values are the accumulated values of the “MCTLSZ” and “MCTLTGT” as described below.



We have to look for the “MCTL” columns to view ballooning activity on a per VM basis:

“MCTL?”: indicates if the balloon driver is active “Y” or not “N”. If VMware tools is not installed or not running this value will show as “N”

“MCTLSZ”: the amount (in MB) of guest physical memory that is actually reclaimed by the balloon driver

“MCTLTGT”: the amount (in MB) of guest physical memory that is going to be reclaimed (targeted memory). If this counter is greater than “MCTLSZ”, the balloon driver inflates causing more memory to be reclaimed. If “MCTLTGT” is less than “MCTLSZ”, then the balloon will deflate. This deflating process runs slowly unless the guest requests memory.

“MCTLMAX”: the maximum amount of guest physical memory that the balloon driver can reclaim. Default is 65% of assigned memory.

Resource Allocation Tab:  

You can verify the Memory Ballooning stats of each individual VM from VM Resource Allocation Tab. This particular VM  Ballooned value is 5.08 GB

  



Transparent Page sharing


The main goal of TPS is provid more memory to VM than physical host has. This is called memory over commitment.

TPS is known as memory deduplication

Transparent page sharing is a method by which redundant copies of pages are eliminated. 



If the hypervisor identifies identical memory pages on multiple virtual machines (VMs) on a host, it shares them among virtual machines (VMs) with pointers. This frees up memory for new pages. If a VM's information on that shared page changes, the hypervisor writes the memory to a new page and readdresses a pointer.

In short ,Identical memory can be shared among VM

How does it work?


The memory is split in 4 KB pages. Windows Virtual machines might have identical memory pages. TPS runs every 60 minute .It scans all memory pages and create a HASH value for each of them.



Those hashes are saved global hash table and then it compared to each other by kernel. Every time the ESX kernel finds two identical hashes the kernel leaves only one copy of page in memory and removes the second one.



when one of your Virtual machine requests to write to the page, VMkernel creates a new page and new page access will only be provided to that particular virtual machine. This terminology is called Copy-on Write (COW).



You can verify the memory which are shared using Transparent memory Sharing (TPS) from Esxtop ,Virtual Machine Resource Allocation tab and also using vCenter Performance Graphs.

esxtop -> Press m

You will be able to see how much % of memory is overcommited in your ESXi host using the Value MEM Overcommit avg. The MEM overcommit avg tells us that the average memory over commitment level averages in 1-min, 5-min and 15-min. A value of 0.50 is a 50% over commitment of memory. In our case it is 5.87 which is nothing but 587% memory over commitment on my host. My ESXi host is having 5 GB of memory with 5 Virtual Machines. Out 5, 4 VM’s are allocated with 8 GB and 1 VM alloacted with 2 GB of memory. My total ESXi memory is 5 Gb but allocated memory for Virtual machines is 34 GB. which is almost 7 times the available memory of my ESXi host. This over commitment becomes only possible because of this VMware Memory management techniques.


Detailed stats about Memory saving using Transparent page sharing can be found with PSHARE value. Take a look at PSHARE/MB 2575 MB which is shared between the Virtual machines out of which 355 MB is common. Which allows us to save 2220 MB of memory using Transparent Page sharing.




Memory which are shared at individual Virtual Machines can also be viewed using the resource allocation tab of each virtual machines. Below Virtual machine is having around shared memory of around 1.64 GB. which is the Amount of guest “physical” memory shared with other virtual machines using the transparent page-sharing mechanism.





You can also use vCenter Performance graphs to collect the Shared memory stats of each Virtual Machine on the ESXi host using the shared stats under Memory in vCenter Advanced chart options.


Shared Common is the Amount of machine memory that is shared by all powered-on virtual machines and vSphere services on the host.

shared – sharedcommon = machine memory (host memory) savings (KB)

2575 Mb – 355MB = 2220 MB  host memory saving





Memory Compression:


In simple the when memory contention happens, the memory pages which is about to swapped will be compressed and store in the main memory instead of disk.

Memory compression outperforms swapping because the data accessed from the main memory than the disk

The compression ratio should always less than 50%

Hypervisor Swapping


In the cases where Ballooning (and TPS) are not sufficient to reclaim memory, ESX employs Hypervisor Swapping to reclaim memory. At guest startup, the hypervisor creates a separate swap file for the guest. This file located in the guest’s home directory has an extension .vswp. Then, if necessary, the hypervisor can directly swap out guest physical memory to that swap file, which frees host physical memory for other guests. The swap file size is set to the guest physical memory minus its Reservation. For example, if you allocated 4GB to a guest and set a Reservation of 1GB, the swap file size will be 3GB

Tuesday, March 31, 2015

vCenter Operations Manager – Major Badges – Health, Risk, and Efficiency



vC Ops has 3 major badges and their status depends on the minor badges whose scores get rolled up to make up the major badges. Those 3 major badges are Health, Risk, and Efficiency. Each of these 3 badges are weighted combinations of minor badges

vCenter Operations Health Badge
health-badge2
The vC Ops Health badge tells you how “healthy” your vSphere infrastructure is (if you are at the World level of the inventory) or it tells you how healthy a particular object is such as a virtual data center, host, VM, or cluster.

The health badge is a weighted combination of Workload, Anomalies and Faults badges.
The higher your health score, the better off you are. Thus, a “100″ is perfect health.

The Health badge summarizes workload, anomalies, and faults.

The Workload badge shows how hard an object is working. A higher workload score indicates that an object is doing more work. Obviously, you don’t want objects out there doing zero work, as that is waste but, as the same time, you also don’t want objects completely maxed out with a workload score of 100 either. Workload is an absolute measurement that calculates the demand for a resource divided by the capacity of an object. Resources might include CPU, memory, disk I/O, or network I/O.  vC Ops will help you to balance workload across your resource objects effectively.

The Anomalies badge indicates how the object is behaving currently compared to how it has behaved in the past. While small anomalies don’t always indicate something bad, large anomalies are likely an indicator of a problem. vC Ops uses anomalies to determine what is “normal” in the your vSphere infrastructure vs what is “abnormal”.

The Faults badge tells you if configuration issues have occurred for an object. Faults are given priority over anomalies and workload when calculating health. Faults are calculated based on the events received from VMware vCenter about an object. Examples of events that might generate faults are ESXi host memory errors, loss of network or HBA redundancy, a failover event in a HA cluster, or hardware events (like high CPU temperature) received from CIM events.

vCenter Operations Risk Badge
risk-badge2
The second major metric that vC Ops report is Risk. Risk is a combination of its three sub-metrics - Stress, Time Remaining and Capacity Remaining. You can think of Risk as a rating of how “risky” the virtual infrastructure is in terms of it’s performance or capacity.The difference is that, with those scores, a lower number is a bad indicator where, with Risk in vC Ops, a higher number is a bad indicator.


With Time Remaining, you will be able to see the amount of time left before the object you are analyzing reaches its maximum capacity.  

The Capacity Remaining badge score indicated the number of remaining virtual machines you can fit in that object. For example, on a datastore the capacity remaining is pretty straightforward — how much capacity is remaining to hold VMs?(the lowest number is the capacity remaining).

Stress badge reports the stress that an object is under. Just as your stress level is related to your workload, so is the stress score in vC Ops.Stress is reported between 0 and 100 with 100 being very high stress and 0 being no stress.


vCenter Operations Efficiency Badge
efficiency-badge2
The third major badge that vC Ops reports is Efficiency.

Efficiency Minor Badges – Reclaimable Waste and Density

 The Reclaimable Waste badge indicates what resources you can get back from your virtual infrastructure. Those reclaimed resources might allow you to provision more VMs.

The Density badge is measures your virtual infrastructure consolidation ratios to  to ensure that you are maximizing your virtual infrastructure investment.

vCenter and ESXi Log Files



vCenter Server Log files
  • vCenter Server 5.x on Windows Server 2003: %ALLUSERSPROFILE%\Application Data\VMware\VMware VirtualCenter\Logs\
  • vCenter Server 5.x on Windows Server 2008: %ALLUSERSPROFILE%\VMware\VMware VirtualCenter\Logs\
  • vCenter Server 5.x Linux Virtual Appliance: /var/log/vmware/vpx/
With a Windows Server implementation of vCenter, browse to the log file location and open the log in your favorite text editor.

ESXi Log Files
  • /var/log/auth.log: ESXi Shell authentication success and failure attempts.
  • /var/log/dhclient.log: DHCP client log.
  • /var/log/esxupdate.log: ESXi patch and update installation logs.
  • /var/log/hostd.log: Host management service logs, including virtual machine and host Task and Events, communication with the vSphere Client and vCenter Server vpxa agent, and SDK connections.
  • /var/log/shell.log: ESXi Shell usage logs, including enable/disable and every command entered.
  • /var/log/boot.gz: A compressed file that contains boot log information and can be read using zcat /var/log/boot.gz|more.
  • /var/log/syslog.log: Management service initialization, watchdogs, scheduled tasks and DCUI use.
  • /var/log/usb.log: USB device arbitration events, such as discovery and pass-through to virtual machines.
  • /var/log/vob.log: VMkernel Observation events, similar to vob.component.event.
  • /var/log/vmkernel.log: Core VMkernel logs, including device discovery, storage and networking device and driver events, and virtual machine startup.
  • /var/log/vmkwarning.log: A summary of Warning and Alert log messages excerpted from the VMkernel logs.
  • /var/log/vmksummary.log: A summary of ESXi host startup and shutdown, and an hourly heartbeat with uptime, number of virtual machines running, and service resource consumption.
Ways to View vSphere Log Files
There are a number of ways in which you can view log files, depending on whether they are for vCenter or for an ESXi host. I’ll start by looking at ways in which you can view host log files. The first place is simply from the DCUI on the host. You can move down to ‘View System Logs’, then choose the log file that you would like to view:

 

The second way is to use the vSphere client. By making a connection directly to a host, rather than vSphere, you can view the hosts log files:
 
 
 

If you are connected to vCenter rather than a host, you can browse to the same place, but instead of host logs, you can view the vCenter logs files.

Another way to view a host’s log files is to use a web browser. This is a method I always forget is available, but it definitely has its uses. Using a url like this one: https://192.168.0.235/host, will (after you have authenticated) present you with a web page from which you can access host log files: