Showing posts with label lvmstat. Show all posts
Showing posts with label lvmstat. Show all posts

Thursday, August 8, 2013

AIX hotspot hunting and potential benefit of tier 0 cache: lvmstat overtime data

Recently I posted a method for tacking on a timestamp to lvmstat output with awk.  Maybe you think that's not a big deal.  In my world, it is :)

The graph above is a heat map of a logical volume with a JFS2 filesystem on top.  Its cio mounted, and has tons of database files in it.  With this particular database engine, all reads and writes to a cio filesystem are 8k (unlike Oracle and SQL Server which will coalesce reads and writes when possible).  Also, all database reads go into database cache, unlike Oracle which can elect to perform direct path reads into PGA rather than into the SGA database cache.  Those considerations contribute to the nice graph above.

The X axis is the logical partition number - the position from beginning to end in the logical volume.
The left primary axis represents the number of minutes in out of 360 elapsed minutes that the logical partition was ranked in the top 32 for the logical volume by iocnt.  (Easy to get with the -c parameter for the lvmstat command.)

So... what use is such a graph?  Together with data from lslv, you can map hot logical partitions back to their physical volumes (LUNs)... that is invaluable when trying to eliminate QFULL occurrences at the physical volume level.  For another thing, together with database cache analysis (insertion rates at all insertion points, expire rates, calculated cache hold time, etc) heat maps like these can help to estimate the value of added database cache.  They can also help to estimate the value of a tier 0 flash automated tiering cache - whether the flash is within the storage array, or in an onboard PCIe flash device.  I've worked a bit with EMC xtremsw for x86 Windows... looking forward to working with it for IBM Power and PCIe flash form factor.  Hoping to test with the QLogic Mt Rainier/FabricCache soon, too.  If the overall IO pattern looks like the graph above, as long as the flash capacity together with the caching algorithm results in good caching for the logical partitions on the upward swing of the hockey stick, you can expect good utilization of the tier 0 cache.  I personally prefer onboard cache, because I like as little traffic getting out of the server as possible.  In part for lower latency... in part to eliminate queuing concerns... but mostly to be a good neighbor to other tenants of shared storage.

So... in reality, the system I'm staring at today isn't as simple as a single logical volume heat map.  There are more than 6 logical volumes that contain persistent database files.  All of those logical volumes have hockey stick shaped graphs as above.  Its easy to count the number of upswing logical partitions across all of those LVs, and find out how much data I really want to see in the flash cache.  Now... if the flash cache transfer size is different than the logical partition size that should be considered.  If my logical partition size is 64mb, and the flash tiering always promotes contiguous 1 GB chunks from each LUN, that could lead to requiring a lot more flash capacity to contain all of my hot data.  On the other hand, if the transfer size into tier 0 cache is 1 mb, and the heat is very uneven among 1 mb chunks within each hot LP... the total flash cache size for huge benefit might be a lot smaller than the aggregate size of all hot LPs.  Something to think about.  But I can't give away all of my secrets. Keeping at least some secrets is key to sasquatch survival :)

Wednesday, July 31, 2013

IBM Power AIX hotspot hunting - About that lvmstat date timestamp...

So a few posts ago I cavalierly wrote how easy it was to add a timestamp to lvmstat for unattended data gathering when hotspot hunting. (Hotspot hunting ain't easy unless you can time-align lvmstat and iostat data.)  I was overconfident, in my own awk abilities at the very least.

Sure, I found the hot logical partitions when I was reviewing data.  But every single one of my timestamps in the log had the same value!  Criminey!  It took me forever to spreadsheet-mod my way out of that in order to create some decent looking spreadsheet and graph data.

If you McGoogle "lvmstat", "awk", and "timestamp" I don't think you'll find very much that is helpful in allowing time-aligning of unattended lvmstat logs with iostat logs.  Maybe there's a compact answer somewhere, but I didn't find it... don't wanna brag but I'm a pretty good McGoogler.  Most of the responses in various forums resulted in the same stale date timestamp values, collected at the beginning of script execution and repeated throughout.  Eventually I was able to get a date timestamp value that was updated throughout execution... but then I had extra linebreaks that I didn't want.


My pain... your gain.  Here's what I ended up with*.

# ###find top 4 partitions, 5 second interval, 4 iterations
# lvmstat -s -l sasquatch_lv -c4 5 4 | awk -u '{ORS=" ";} {print $0 ; system("date") ; close("date") ; }'
 Wed Jul 31 16:40:16 CDT 2013
Log_part  mirror#  iocnt   Kb_read   Kb_wrtn      Kbps Wed Jul 31 16:40:16 CDT 2013
      45       1   39773        16    340312      0.02 Wed Jul 31 16:40:17 CDT 2013
       1       1   16298         0    425208      0.02 Wed Jul 31 16:40:17 CDT 2013
       3       1    5428         0    685612      0.03 Wed Jul 31 16:40:17 CDT 2013
       2       1    4528         0    480760      0.02 Wed Jul 31 16:40:17 CDT 2013
... Wed Jul 31 16:40:26 CDT 2013


Yay!  A small victory, but I'll take it.

*This isn't really the end of course... there's a bit more cleanup.  The empty first line can be removed, the lines that represent no change in values and have only periods and timestamps can be removed.  Those are left as exercises for the reader.  :)


***** Update sql_sasquatch 10/03/2013 *****
I pulled this out today to make sure it really works and that I wasn't just conveniently ignoring timestamps that weren't updating as desired.  Nope, its working like I thought. 
# lvmstat -s -l lv_unicorn -c4 5 20 | awk -u '{ORS=" ";} {print $0 ; system("date") ; close("date") ; }'
 Thu Oct  3 15:00:45 CDT 2013
Log_part  mirror#  iocnt   Kb_read   Kb_wrtn      Kbps Thu Oct  3 15:00:45 CDT 2013
      77       1     146         0       592      0.06 Thu Oct  3 15:00:45 CDT 2013
      75       1      91       144       240      0.04 Thu Oct  3 15:00:45 CDT 2013
      78       1      88         4       348      0.04 Thu Oct  3 15:00:45 CDT 2013
      85       1      80         4      2412      0.26 Thu Oct  3 15:00:45 CDT 2013
. Thu Oct  3 15:00:55 CDT 2013
      85       1       2         0        52     11.11 Thu Oct  3 15:00:55 CDT 2013
....... Thu Oct  3 15:01:35 CDT 2013
      75       1       7         0        28      5.80 Thu Oct  3 15:01:35 CDT 2013
      85       1       7         0       148     30.67 Thu Oct  3 15:01:35 CDT 2013
      77       1       5         0        20      4.15 Thu Oct  3 15:01:35 CDT 2013
      40       1       4         0        16      3.32 Thu Oct  3 15:01:35 CDT 2013
........ Thu Oct  3 15:02:15 CDT 2013
      85       1       2         0       124     21.96 Thu Oct  3 15:02:15 CDT 2013

Monday, July 22, 2013

IBMPower AIX Hotspot hunting - lvmstat

So, maybe you think there are hotspots in disk IO queuing.  I often think that when looking at #Oracle 11GR2 on #AIX #IBMPower servers.  The lvmstat command can be your best friend in tracking down host side LVM hotspots.  Believe me, sometimes you can get a big benefit by moving a small amount of data.  But, usually you end up proving some known best practices - like keeping Oracle redo logs on separate logical and physical volumes from Oracle database .dbf files, and keeping ETL flat files on separate logical/physical volumes from both redo logs and dbf database files :)

At any rate, lvmstat is a mega-useful tool.    A few references for your reading enjoyment.
http://pic.dhe.ibm.com/infocenter/aix/v6r1/topic/com.ibm.aix.prftungd/doc/prftungd/lvm_perf_mon_lvmstat.htm
http://poweritpro.com/performance/if-your-disks-are-busy-call-lvmstat

Root privileges are required for lvmstat, and unlike iostat there isn't a parameter to timestamp its output.  Easily remedied with awk.


Here's some iostat info from mountainhome, a server chosen totally at random. :)
# iostat -DlRTV hdisk0

System configuration: lcpu=24 drives=6 paths=10 vdisks=2

Disks:                     xfers                                read                                write                                  queue                    time
-------------- -------------------------------- ------------------------------------ ------------------------------------ -------------------------------------- ---------
                 %tm    bps   tps  bread  bwrtn   rps    avg    min    max time fail   wps    avg    min    max time fail    avg    min    max   avg   avg  serv
                 act                                    serv   serv   serv outs              serv   serv   serv outs        time   time   time  wqsz  sqsz qfull
hdisk0           0.6  44.3K   6.3  11.7K  32.6K   2.5   5.3    0.1  215.6     0    0   3.8   1.8    0.2  165.3     0    0  11.8    0.0  130.5    0.0   0.0   2.8  12:10:08

My eagle-sharp eyes train in on the serv qfull number - a rate of 2.8 per second for the monitoring period!  Time for intervention!  At least for me - I hate qfulls.

Lets try to be scientifical about this.


Enable the volume groups/logical volumes for stats collection.
A script I've got for that - ain't pretty but it works. (execute as root)
#enable lvmstat stats
for vg in `lsvg | /usr/bin/awk -u '{print $1 ;}'`;
   do
   lvmstat -v $vg -e
   for lv in `lsvg -l $vg | /usr/bin/awk -u 'NR <= 2 { next } {print $1 ;}'`;
      do
         lvmstat -l $lv -e ;
      done ;
   done ;


Once stats are enabled for the volume groups/logical volumes that you care about... start collecting!  (execute as root)
./sasquatch_lvmstat_script 3 3 $(pwd) 4 &


#sasquatch_lvmstat_script
# Begin Functions
lvmon_cmd(){
/usr/sbin/lvmstat -s -l $1 -c$2 $3 $4 | /usr/bin/awk -u '{print $0, date}' date="`date '+%D %H:%M:%S'`" >> $5/$(hostname)_$6_$1_$(date +%Y%m%d)
}
# End functions

interval=${1:-2}
count=${2:-2}
logdir=${3:-/tmp}
top=${4:-32}
for vg in `lsvg | /usr/bin/awk -u '{print $1 ;}'`;
   do
   for lv in `lsvg -l $vg | /usr/bin/awk -u 'NR <= 2 { next } {print $1 ;}'`;
      do
         lvmon_cmd $lv $top $interval $count $logdir $vg $lv &
      done ;
   done ;


Lets pretend I've looked at enough of the output files to determine that lvmcore is the heavy hitter logical volume in rootvg.

The output files from the lvmstat script look kinda like this:

cat mountainhome_rootvg_lvcore_20130722
 07/22/13 11:42:47
Log_part  mirror#  iocnt   Kb_read   Kb_wrtn      Kbps 07/22/13 11:42:47
      56       1      53       212         0      0.00 07/22/13 11:42:47
       1       1      52       208         0      0.00 07/22/13 11:42:47
      53       1      10        40         0      0.00 07/22/13 11:42:47
       2       1       7        28         0      0.00 07/22/13 11:42:47
.. 07/22/13 11:42:47



Log_part are the logical partitions of the logical volume.  Match them up to physical partitions and physical volumes based on the output of something like 'lslv -m'.
# lslv -m lvcore | grep -e 0056 -e 0053 -e 0001 -e 0002
0001  0704 hdisk0
0002  0705 hdisk0
0053  0756 hdisk0
0056  0759 hdisk0

So... what are the lines with nothing but periods and a timestamp?  Those are quiet collection intervals - lvmstat collapsed them for you.  That was nice!  Kinda like including -V in iostat -DlRTV you only get what you need.

OK.  So Logical partitions 0056 and 0001 are the busiest in logical volume lvcore.  They are both on hdisk0.  And we supposedly compared the output across our logical volumes in rootvg and saw that, while hdisk0 is under qfull pressure, lvmcore is the heavy hitter in rootvg volume group.  If there's a quiet hdisk in rootvg, moving one of those logical partitions - or both of them... could relieve the pressure.  Or, maybe you know the contents.  (The fileplace command can assist by mapping individual file fragments to logical partitions of the logical volume - that's a lesson for another day.)  Sometimes its easiest to say "golly!  I guess I could put such-n-such into a different filesystem on a different volume group!  It wouldn't put any pressure on filesystem buffers or any pressure on those physical volume IO queues then!"  That kinda stuff requires negotiation, though.  And I'm not a negotiator, just a sasquatch.

When you are done with lvmstat monitoring, disable the stats collection.  Its a small amount of overhead, but if you aren't baselining or investigating... no need to have the stats enabled. (execute as root)

#disable lvmstat stats
for vg in `lsvg | /usr/bin/awk -u '{print $1 ;}'`;
   do
   /usr/sbin/lvmstat -v $vg -d
   for lv in `lsvg -l $vg | /usr/bin/awk -u 'NR <= 2 { next } {print $1 ;}'`;
      do
         /usr/sbin/lvmstat -l $lv -d ;
      done ;
   done ;