Other Can IOSTAT be trusted ? to report actual disk write

Hi
Trying to estimate how long a SSD device will live before failing .
if I have a Kingston A400 (cheap consumer grade ) 240GB SSD with a TBW value of 80 Terabyte.

80.000 / ( 365 x 3 x240 ) = 0.3 DiskWrites per day
240* 0,3 = 73

so for three years the A400 is listed to be able to receive 73 Gbytes per day.
and for five years the A400 is listed to be able to receive 44 Gbytes per day.

is there a known method to measure how much data is written to and deleted from a storage device during a delta T. ?
will IOSTAT tell me the truth on amount of Megabytes written to the medium during the interval set ?
Does IOSTAT include all data written to /dev/nda0 ? /dev/nda1 ......
Swap data ? UFS journal ? ZFS LOG & CACHE ?

what traffic to the storage device is not measured by IOSTAT ?
 
iostat will report writes as they happen through the VFS subsystem. That means it can count too many writes when some writes are superseded by new writes to the same location and never make it out of the VM cache.

Good SSDs have a direct counter for blocks written.
 
SMART Attribute ID 241 "Lifetime Writes from Host System" will give you total GB written on the disk and ID 177 will give you the wear levels
 
232 Available_Reservd_Space 0x0033 100 100 004 Pre-fail Always - 100
233 Media_Wearout_Indicator 0x0032 100 100 --- Old_age Always - 310
These are the values I'd look at on an SSD. No worries here, as 100 is still a long ways from 004.
 
iostat will report writes as they happen through the VFS subsystem. That means it can count too many writes when some writes are superseded by new writes to the same location and never make it out of the VM cache.
Not sure if this is true. I think that iostat takes stats from devstat(9) and that takes input at the the storage layer (e.g., commonly CAM).
 
Late to the party here, but in the interest of sharing what I've learned I'm leaving this for posterity.

I have two Kingston KC600 256 GB drives that I bought last year, right before prices exploded. Supposed to be a decent enough consumer SSD. They're set up in a zfs mirror for zroot. I run a mix of 17 jails and 3 bhyve VM's on my server running a bunch of podman containers. I set up prometheus and grafana in a jail to have some basic metrics, and I see the S.M.A.R.T indicator for remaining lifetime ticking down with about 1% per week, mostly in sync for both drives.

1784668590763.png


Here's one of them:
=== START OF INFORMATION SECTION ===
Model Family: Silicon Motion based SSDs
Device Model: KINGSTON SKC600256G
Serial Number: --- redacted ---
LU WWN Device Id: 5 0026b7 687536fa5
Firmware Version: S4800116
User Capacity: 256,060,514,304 bytes [256 GB]
Sector Sizes: 512 bytes logical, 4096 bytes physical
Rotation Rate: Solid State Device
Form Factor: 2.5 inches
TRIM Command: Available, deterministic, zeroed
Device is: In smartctl database 7.5/5706
ATA Version is: ACS-3 T13/2161-D revision 5
SATA Version is: SATA 3.3, 6.0 Gb/s (current: 6.0 Gb/s)
Local Time is: Tue Jul 21 22:54:44 2026 CEST
SMART support is: Available - device has SMART capability.
SMART support is: Enabled

=== START OF READ SMART DATA SECTION ===
SMART overall-health self-assessment test result: PASSED

General SMART Values:
Offline data collection status: (0x80) Offline data collection activity
was never started.
Auto Offline Data Collection: Enabled.
Self-test execution status: ( 0) The previous self-test routine completed
without error or no self-test has ever
been run.
Total time to complete Offline
data collection: ( 0) seconds.
Offline data collection
capabilities: (0x7b) SMART execute Offline immediate.
Auto Offline data collection on/off support.
Suspend Offline collection upon new
command.
Offline surface scan supported.
Self-test supported.
Conveyance Self-test supported.
Selective Self-test supported.
SMART capabilities: (0x0002) Does not save SMART data before
entering power-saving mode.
Supports SMART auto save timer.
Error logging capability: (0x01) Error logging supported.
General Purpose Logging supported.
Short self-test routine
recommended polling time: ( 2) minutes.
Extended self-test routine
recommended polling time: ( 30) minutes.
Conveyance self-test routine
recommended polling time: ( 2) minutes.
SCT capabilities: (0x0031) SCT Status supported.
SCT Feature Control supported.
SCT Data Table supported.

SMART Attributes Data Structure revision number: 16
Vendor Specific SMART Attributes with Thresholds:
ID# ATTRIBUTE_NAME FLAG VALUE WORST THRESH TYPE UPDATED WHEN_FAILED RAW_VALUE
1 Raw_Read_Error_Rate 0x0000 100 100 000 Old_age Offline - 0
5 Reallocated_Sector_Ct 0x0000 100 100 000 Old_age Offline - 0
9 Power_On_Hours 0x0000 100 100 000 Old_age Offline - 6312
12 Power_Cycle_Count 0x0000 100 100 000 Old_age Offline - 46
148 Total_SLC_Erase_Ct 0x0000 100 100 000 Old_age Offline - 13051
149 Max_SLC_Erase_Ct 0x0000 100 100 000 Old_age Offline - 887
150 Min_SLC_Erase_Ct 0x0000 100 100 000 Old_age Offline - 633
151 Average_SLC_Erase_Ct 0x0000 100 100 000 Old_age Offline - 870
159 DRAM_1_Bit_Error_Count 0x0000 100 100 000 Old_age Offline - 0
160 Uncorrectable_Error_Cnt 0x0000 100 100 000 Old_age Offline - 0
161 Valid_Spare_Block_Cnt 0x0000 100 100 000 Old_age Offline - 40
164 Total_Erase_Count 0x0000 100 100 000 Old_age Offline - 283060
165 Max_Erase_Count 0x0000 100 100 000 Old_age Offline - 546
166 Min_Erase_Count 0x0000 100 100 000 Old_age Offline - 446
167 Average_Erase_Count 0x0000 100 100 000 Old_age Offline - 528
169 Remaining_Lifetime_Perc 0x0000 056 056 000 Old_age Offline - 56
177 Wear_Leveling_Count 0x0000 100 100 050 Old_age Offline - 36317
181 Program_Fail_Cnt_Total 0x0000 100 100 000 Old_age Offline - 0
182 Erase_Fail_Count_Total 0x0000 100 100 000 Old_age Offline - 0
192 Power-Off_Retract_Count 0x0000 100 100 000 Old_age Offline - 9
194 Temperature_Celsius 0x0000 041 053 000 Old_age Offline - 41
195 Hardware_ECC_Recovered 0x0000 100 100 000 Old_age Offline - 0
196 Reallocated_Event_Count 0x0000 100 100 016 Old_age Offline - 0
199 UDMA_CRC_Error_Count 0x0000 100 100 050 Old_age Offline - 0
231 SSD_Life_Left 0x0000 056 056 000 Old_age Offline - 56
232 Available_Reservd_Space 0x0000 100 100 000 Old_age Offline - 100
241 Host_Writes_32MiB 0x0000 100 100 000 Old_age Offline - 809347
242 Host_Reads_32MiB 0x0000 100 100 000 Old_age Offline - 54402
245 TLC_Writes_32MiB 0x0000 100 100 000 Old_age Offline - 4670490

SMART Error Log Version: 1
No Errors Logged

This drive has currently been powered on for about 6300 hours, or 252 days or so. It has written 241 Host_Writes_32MiB = 809347 in 32 MiB units, which equals about 25 TB of data. But it also logs 245 TLC_Writes_32MiB = 4670490 or around 153TB, which if I understand this correctly is actual NAND flash writes. That is close to 6 x write amplification. Looking at attributes 148-161, there's a lot of internal movement between SLC cache and TLC storage. Those count towards attribute 245, even though they're not host writes.If the 528 erase cycles scale linearly and I have 44% "used" lifetime, that means the drive is rated for approximately 1200 cycles. For a 256 GB SSD × 1200 cycles equals about 307 TB of NAND writes, and given this disk so far has about 153 TB: 153 / 307 ≈ 50% use. So that tracks roughly with the lifetime remaining attribute.

The datasheet for these drives says the 256 GB version is rated for 150TBW, which is more than I derive from attribute 241, and less than I derive from attribute 245, seen against the "remaining lifetime percentage". I'm a litte unsure about that. For what it's worth I think the write amplification is larger than expected in a system like mine, an I think that rating is probably closer to the estimate for erase cycles (that I estimate at 1200).

What I've taken from this is that my two Kingston SSDs will reach the end of their useful life after ~500 days in service, which is way less than I expected when I bought them. Consumer SSDs in a write-amplified homeserver seeing 24/7 use with ZFS is a bit optimistic. If I got these in the 1TB size, they would likely last longer, and of course, the drives might work well beyond that "counter" but I'm not inclined to trust them in a server far beyond manufacturer's estimate. So I've bought some used enterprise SSD's I will replace them with in a few months.
 
Note that some older SSD drives do not report the remaining life as an attribute but show it differently. In my case I have a pretty good Toshiba SSD drive I bought at Fry's Electronics from 11-12 years ago, and the only way I can see the life left percentage is by running smartctl -x /dev/adaN. It shows that under "Device statistics" section!

Why did Fry's Electronics go out of business anyway? So sad! It was like a chain of giant electronics store in Cali that seemed to be doing really well. They themed their stores too - one would have a tropical theme with palm trees and a safari or whatever, another something else.

I did not miss when Circuit City went out of business, though. I thought Best Buy would end up that way too.
 
To really decode SMART information correctly, you need very detailed information from the drive. I used to do this kind of stuff (about 10 years ago, for a large file/storage system), and we had detailed interface manuals for the drives (under NDA), and regular direct contact with their engineers. We would literally write e-mails to our engineer friends at Seagate/WD/Sandisk/Sandforce/... a few times a month, have phone meetings, and get good answers. Obviously, this approach won't work on (a) SSDs or hard disks intended for the consumer market, and (b) individual end user customers. You can try to see whether full interface manuals are available for these drives; Seagate often publishes them (to anyone) for their drives, but I don't know whether the SMART data is fully documented there, or only the SCSI/SATA command set.

Doing this right is a heck of a lot of work. In a business setting, I would assign an engineer for a month or a quarter.

Why did Fry's Electronics go out of business anyway? So sad!
Combination of things. First, the whole business model of retailing white goods (housewares), electronics (stereos...), computers, and electronics components fell apart slowly between when they started in the 80s, and when they shut down in the last few years. A lot of consumers didn't want to shop there, since the store was so ridiculously badly run (disorganized, lots of returned goods rewrapped and back on the shelves, the folks who worked there had no clue and a bad attitude). They were famously obnoxious about refusing returns ... and when they did take returns, they would put them back on the shelves. But the real problem was top management. They had a lot of money to start with, since the Fry family had made a fortune on a grocery store chain. They invested it stupidly, into buying too many stores (including the former "Incredible Universe" stores), and putting too much money into making them look fancy.

One of the reasons I liked them: They used to stock chips, electronic parts (breadboard, soldering stuff), resistors, transistors, cases. You could be an electronics hobbyist using them as the main supply depot. But: The number of parts they needed to stock was HUGE (the 74xxx series of chips alone is a thousand different things, if you count 74, LS, HCT ... varieties), so they were losing money on having too much stuff with low profit margins. It was insanely labor intensive for them. The moment internet ordering took off, and distributors (such as Anthem/Avnet/Digikey/Mouser/Newark/... ) became consumer-friendly, that whole business collapsed, and they didn't figure it out.

And to a large extent, that was the story of their bad management: They didn't figure it out when the world changed. Amazon, Best Buy, Digikey, and computer stores all ate Fry's lunch, while the owners/executives were absent or clueless.

Add to that a whole slew of outright theft. One of their VPs (I think of purchasing) was a massive gambling addict, who stole about $150M from Fry's to feed his gambling habit. The court cases made the news all over. That's not terribly much money for a large retail chain, but it shows that management wasn't watching anything and not even noticing losses. Complete incompetence in the C-suite.
 
I never had a problem returning stuff there - not that I did a lot of it but a couple of times I had to. They had a little bit of an attitude in general but beats having great customer service while ripping you off. But sometimes you had to actively search for one of their guys to actually even purchase some part. It was really annoying that you had to fetch someone to print you a sales order or whatever before you could go and pay at the checkout.

I am with you on loving them for selling all sorts of chips and electronics parts. You could see how the inventory shuffling could've eaten into their low profit margins. Because they did keep the prices low.

I dunno, you're saying all those other businesses ate Fry's lunch, but I feel Fry's had a very solid separate market and loyal customer base. But what you are saying has some merit in terms of online shopping. They had a horror show of a website (the domain was outpost.com? like wth) that was as clunky as a hand cranked engine from a century ago. Really unacceptable - possibly due to their specific business model focusing on stores, but it'd be such an easy thing for them to add to their operations - that would probably allow them to thrive.
 
Back
Top