[{"data":1,"prerenderedAt":3254},["ShallowReactive",2],{"doc:\u002Fadvanced-data-transformation-and-cleaning\u002Fhandling-missing-data-in-excel-reports\u002Ffind-and-report-missing-values-in-an-excel-file":3,"surround:\u002Fadvanced-data-transformation-and-cleaning\u002Fhandling-missing-data-in-excel-reports\u002Ffind-and-report-missing-values-in-an-excel-file":3246},{"id":4,"title":5,"body":6,"dateModified":3223,"datePublished":3223,"description":3224,"extension":3225,"faq":3226,"meta":3237,"navigation":213,"path":3238,"seo":3239,"slug":3242,"stem":3243,"type":3244,"__hash__":3245},"docs\u002Fadvanced-data-transformation-and-cleaning\u002Fhandling-missing-data-in-excel-reports\u002Ffind-and-report-missing-values-in-an-excel-file\u002Findex.md","Find and Report Missing Values in an Excel File",{"type":7,"value":8,"toc":3210},"minimark",[9,28,133,138,169,172,429,433,439,599,607,611,614,1039,1042,1049,1124,1128,1131,1247,1250,1486,1493,1497,1500,1607,1868,1871,1875,1991,1994,2735,2742,2746,2858,2862,2868,2879,2907,2913,2916,3066,3074,3078,3087,3091,3126,3144,3150,3156,3166,3170,3206],[10,11,12,13,17,18,21,22,27],"p",{},"Before you fill a gap you have to know it is there, how big it is, and whether it is a gap at all. An Excel file arriving from elsewhere hides missing data in several forms: genuinely empty cells, the string ",[14,15,16],"code",{},"n\u002Fa"," that pandas reads as ordinary text, spacer rows that inflate every count, and sentinel numbers like ",[14,19,20],{},"-1"," that look like data. This guide builds a completeness audit that finds all of them and produces a report somebody can act on. It is the first step in ",[23,24,26],"a",{"href":25},"\u002Fadvanced-data-transformation-and-cleaning\u002Fhandling-missing-data-in-excel-reports\u002F","Handling Missing Data in Excel Reports",".",[29,30,39,40,39,44,39,48,39,55,39,65,39,72,39,77,39,82,39,87,39,92,39,95,39,99,39,104,39,109,39,113,39,118,39,122,39,126,39,129],"svg",{"viewBox":31,"role":32,"ariaLabel":33,"ariaLabelledBy":34,"xmlns":37,"style":38},"0 0 800 254","img","Four forms missing data takes in an Excel file: truly empty cells, disguised text sentinels, whole blank spacer rows, and numeric sentinels like minus one.",[35,36],"miss-t","miss-d","http:\u002F\u002Fwww.w3.org\u002F2000\u002Fsvg","width:100%;max-width:800px;height:auto;display:block;margin:1.5rem auto;font-family:Inter,ui-sans-serif,system-ui,sans-serif","\n  ",[41,42,43],"title",{"id":35},"The four disguises missing data wears",[45,46,47],"desc",{"id":36},"Four categories. Truly empty cells are recognised by pandas as NaN and counted automatically. Text sentinels such as n slash a, a dash or the word unknown are read as ordinary strings and are invisible to isna. Whole blank spacer rows come from the source layout rather than from records and inflate every per-column count. Numeric sentinels like minus one or nine nines look like valid data and quietly distort averages and sums.",[49,50],"rect",{"x":51,"y":51,"width":52,"height":53,"fill":54},"0","800","254","#ffffff",[49,56],{"x":57,"y":58,"width":59,"height":60,"rx":61,"fill":62,"stroke":63,"style":64},"14","26","380","98","13","#d9f4f1","var(--teal,#0f9488)","stroke-width:2px",[66,67,71],"text",{"x":68,"y":69,"style":70},"204","52","font-size:12.5px;font-weight:700;fill:var(--teal-ink,#0b6157);text-anchor:middle","1 · truly empty cells",[66,73,76],{"x":68,"y":74,"style":75},"78","font-size:11px;fill:var(--text,#172033);text-anchor:middle","pandas reads them as NaN",[66,78,81],{"x":68,"y":79,"style":80},"102","font-size:11px;font-weight:700;fill:var(--teal-ink,#0b6157);text-anchor:middle","isna() finds these — the easy case",[49,83],{"x":84,"y":58,"width":59,"height":60,"rx":61,"fill":85,"stroke":86,"style":64},"406","#fee8f2","var(--accent,#f43f8f)",[66,88,91],{"x":89,"y":69,"style":90},"596","font-size:12.5px;font-weight:700;fill:var(--accent-ink,#be185d);text-anchor:middle","2 · text sentinels",[66,93,94],{"x":89,"y":74,"style":75},"\"n\u002Fa\" · \"-\" · \"unknown\" · \" \"",[66,96,98],{"x":89,"y":79,"style":97},"font-size:11px;font-weight:700;fill:var(--accent-ink,#be185d);text-anchor:middle","read as ordinary text — invisible",[49,100],{"x":57,"y":101,"width":59,"height":60,"rx":61,"fill":102,"stroke":103,"style":64},"136","#fdefd8","var(--gold,#b4740a)",[66,105,108],{"x":68,"y":106,"style":107},"162","font-size:12.5px;font-weight:700;fill:var(--gold-ink,#7a4e06);text-anchor:middle","3 · whole blank rows",[66,110,112],{"x":68,"y":111,"style":75},"188","spacers from the sheet layout",[66,114,117],{"x":68,"y":115,"style":116},"212","font-size:11px;font-weight:700;fill:var(--gold-ink,#7a4e06);text-anchor:middle","inflate every per-column count",[49,119],{"x":84,"y":101,"width":59,"height":60,"rx":61,"fill":120,"stroke":121,"style":64},"#ebebfd","var(--brand,#5b5cf0)",[66,123,125],{"x":89,"y":106,"style":124},"font-size:12.5px;font-weight:700;fill:var(--brand-strong,#4338ca);text-anchor:middle","4 · numeric sentinels",[66,127,128],{"x":89,"y":111,"style":75},"-1 · 0 · 999999 · 1900-01-01",[66,130,132],{"x":89,"y":115,"style":131},"font-size:11px;font-weight:700;fill:var(--brand-strong,#4338ca);text-anchor:middle","look valid — distort every average",[134,135,137],"h2",{"id":136},"prerequisites","Prerequisites",[139,140,145],"pre",{"className":141,"code":142,"language":143,"meta":144,"style":144},"language-bash shiki shiki-themes github-light github-dark-high-contrast","pip install pandas openpyxl xlsxwriter\n","bash","",[14,146,147],{"__ignoreMap":144},[148,149,152,156,160,163,166],"span",{"class":150,"line":151},"line",1,[148,153,155],{"class":154},"sMTad","pip",[148,157,159],{"class":158},"srMev"," install",[148,161,162],{"class":158}," pandas",[148,164,165],{"class":158}," openpyxl",[148,167,168],{"class":158}," xlsxwriter\n",[10,170,171],{},"A file containing all four disguises:",[139,173,177],{"className":174,"code":175,"language":176,"meta":144,"style":144},"language-python shiki shiki-themes github-light github-dark-high-contrast","import numpy as np\nimport pandas as pd\n\npd.DataFrame({\n    \"order_id\": [1001, 1002, 1003, None, 1005, 1006],\n    \"region\": [\"North\", \"n\u002Fa\", \"West\", None, \"-\", \"South\"],\n    \"units\": [120, 88, None, None, 95, -1],\n    \"revenue\": [5150.00, 4268.50, 3511.25, None, np.nan, 1820.00],\n    \"note\": [None, None, None, None, \"restated\", None],\n}).to_excel(\"orders.xlsx\", index=False)\n","python",[14,178,179,195,208,215,221,263,300,339,372,405],{"__ignoreMap":144},[148,180,181,185,189,192],{"class":150,"line":151},[148,182,184],{"class":183},"s-kum","import",[148,186,188],{"class":187},"skGVy"," numpy ",[148,190,191],{"class":183},"as",[148,193,194],{"class":187}," np\n",[148,196,198,200,203,205],{"class":150,"line":197},2,[148,199,184],{"class":183},[148,201,202],{"class":187}," pandas ",[148,204,191],{"class":183},[148,206,207],{"class":187}," pd\n",[148,209,211],{"class":150,"line":210},3,[148,212,214],{"emptyLinePlaceholder":213},true,"\n",[148,216,218],{"class":150,"line":217},4,[148,219,220],{"class":187},"pd.DataFrame({\n",[148,222,224,227,230,234,237,240,242,245,247,250,252,255,257,260],{"class":150,"line":223},5,[148,225,226],{"class":158},"    \"order_id\"",[148,228,229],{"class":187},": [",[148,231,233],{"class":232},"sP0c6","1001",[148,235,236],{"class":187},", ",[148,238,239],{"class":232},"1002",[148,241,236],{"class":187},[148,243,244],{"class":232},"1003",[148,246,236],{"class":187},[148,248,249],{"class":232},"None",[148,251,236],{"class":187},[148,253,254],{"class":232},"1005",[148,256,236],{"class":187},[148,258,259],{"class":232},"1006",[148,261,262],{"class":187},"],\n",[148,264,266,269,271,274,276,279,281,284,286,288,290,293,295,298],{"class":150,"line":265},6,[148,267,268],{"class":158},"    \"region\"",[148,270,229],{"class":187},[148,272,273],{"class":158},"\"North\"",[148,275,236],{"class":187},[148,277,278],{"class":158},"\"n\u002Fa\"",[148,280,236],{"class":187},[148,282,283],{"class":158},"\"West\"",[148,285,236],{"class":187},[148,287,249],{"class":232},[148,289,236],{"class":187},[148,291,292],{"class":158},"\"-\"",[148,294,236],{"class":187},[148,296,297],{"class":158},"\"South\"",[148,299,262],{"class":187},[148,301,303,306,308,311,313,316,318,320,322,324,326,329,331,334,337],{"class":150,"line":302},7,[148,304,305],{"class":158},"    \"units\"",[148,307,229],{"class":187},[148,309,310],{"class":232},"120",[148,312,236],{"class":187},[148,314,315],{"class":232},"88",[148,317,236],{"class":187},[148,319,249],{"class":232},[148,321,236],{"class":187},[148,323,249],{"class":232},[148,325,236],{"class":187},[148,327,328],{"class":232},"95",[148,330,236],{"class":187},[148,332,333],{"class":183},"-",[148,335,336],{"class":232},"1",[148,338,262],{"class":187},[148,340,342,345,347,350,352,355,357,360,362,364,367,370],{"class":150,"line":341},8,[148,343,344],{"class":158},"    \"revenue\"",[148,346,229],{"class":187},[148,348,349],{"class":232},"5150.00",[148,351,236],{"class":187},[148,353,354],{"class":232},"4268.50",[148,356,236],{"class":187},[148,358,359],{"class":232},"3511.25",[148,361,236],{"class":187},[148,363,249],{"class":232},[148,365,366],{"class":187},", np.nan, ",[148,368,369],{"class":232},"1820.00",[148,371,262],{"class":187},[148,373,375,378,380,382,384,386,388,390,392,394,396,399,401,403],{"class":150,"line":374},9,[148,376,377],{"class":158},"    \"note\"",[148,379,229],{"class":187},[148,381,249],{"class":232},[148,383,236],{"class":187},[148,385,249],{"class":232},[148,387,236],{"class":187},[148,389,249],{"class":232},[148,391,236],{"class":187},[148,393,249],{"class":232},[148,395,236],{"class":187},[148,397,398],{"class":158},"\"restated\"",[148,400,236],{"class":187},[148,402,249],{"class":232},[148,404,262],{"class":187},[148,406,408,411,414,416,420,423,426],{"class":150,"line":407},10,[148,409,410],{"class":187},"}).to_excel(",[148,412,413],{"class":158},"\"orders.xlsx\"",[148,415,236],{"class":187},[148,417,419],{"class":418},"sa561","index",[148,421,422],{"class":183},"=",[148,424,425],{"class":232},"False",[148,427,428],{"class":187},")\n",[134,430,432],{"id":431},"step-1-count-what-pandas-already-sees","Step 1 — Count what pandas already sees",[10,434,435,438],{},[14,436,437],{},"isna"," covers the genuinely empty cells:",[139,440,442],{"className":174,"code":441,"language":176,"meta":144,"style":144},"import pandas as pd\n\ndf = pd.read_excel(\"orders.xlsx\")\n\nmissing = df.isna().sum()\ncompleteness = (1 - df.isna().mean()) * 100\n\nreport = pd.DataFrame({\n    \"missing\": missing,\n    \"present\": len(df) - missing,\n    \"complete_pct\": completeness.round(1),\n}).sort_values(\"missing\", ascending=False)\n\nprint(report.to_string())\n",[14,443,444,454,458,472,476,486,510,514,524,532,551,565,585,590],{"__ignoreMap":144},[148,445,446,448,450,452],{"class":150,"line":151},[148,447,184],{"class":183},[148,449,202],{"class":187},[148,451,191],{"class":183},[148,453,207],{"class":187},[148,455,456],{"class":150,"line":197},[148,457,214],{"emptyLinePlaceholder":213},[148,459,460,463,465,468,470],{"class":150,"line":210},[148,461,462],{"class":187},"df ",[148,464,422],{"class":183},[148,466,467],{"class":187}," pd.read_excel(",[148,469,413],{"class":158},[148,471,428],{"class":187},[148,473,474],{"class":150,"line":217},[148,475,214],{"emptyLinePlaceholder":213},[148,477,478,481,483],{"class":150,"line":223},[148,479,480],{"class":187},"missing ",[148,482,422],{"class":183},[148,484,485],{"class":187}," df.isna().sum()\n",[148,487,488,491,493,496,498,501,504,507],{"class":150,"line":265},[148,489,490],{"class":187},"completeness ",[148,492,422],{"class":183},[148,494,495],{"class":187}," (",[148,497,336],{"class":232},[148,499,500],{"class":183}," -",[148,502,503],{"class":187}," df.isna().mean()) ",[148,505,506],{"class":183},"*",[148,508,509],{"class":232}," 100\n",[148,511,512],{"class":150,"line":302},[148,513,214],{"emptyLinePlaceholder":213},[148,515,516,519,521],{"class":150,"line":341},[148,517,518],{"class":187},"report ",[148,520,422],{"class":183},[148,522,523],{"class":187}," pd.DataFrame({\n",[148,525,526,529],{"class":150,"line":374},[148,527,528],{"class":158},"    \"missing\"",[148,530,531],{"class":187},": missing,\n",[148,533,534,537,540,543,546,548],{"class":150,"line":407},[148,535,536],{"class":158},"    \"present\"",[148,538,539],{"class":187},": ",[148,541,542],{"class":232},"len",[148,544,545],{"class":187},"(df) ",[148,547,333],{"class":183},[148,549,550],{"class":187}," missing,\n",[148,552,554,557,560,562],{"class":150,"line":553},11,[148,555,556],{"class":158},"    \"complete_pct\"",[148,558,559],{"class":187},": completeness.round(",[148,561,336],{"class":232},[148,563,564],{"class":187},"),\n",[148,566,568,571,574,576,579,581,583],{"class":150,"line":567},12,[148,569,570],{"class":187},"}).sort_values(",[148,572,573],{"class":158},"\"missing\"",[148,575,236],{"class":187},[148,577,578],{"class":418},"ascending",[148,580,422],{"class":183},[148,582,425],{"class":232},[148,584,428],{"class":187},[148,586,588],{"class":150,"line":587},13,[148,589,214],{"emptyLinePlaceholder":213},[148,591,593,596],{"class":150,"line":592},14,[148,594,595],{"class":232},"print",[148,597,598],{"class":187},"(report.to_string())\n",[10,600,601,602,236,604,606],{},"That is the baseline, and it understates the problem on any real file — ",[14,603,16],{},[14,605,20],{}," and the spacer row are all counted as present.",[134,608,610],{"id":609},"step-2-catch-the-disguised-blanks","Step 2 — Catch the disguised blanks",[10,612,613],{},"Sentinels are file-specific, so make the list explicit rather than hoping the defaults cover it:",[139,615,617],{"className":174,"code":616,"language":176,"meta":144,"style":144},"import pandas as pd\n\nTEXT_SENTINELS = {\n    \"\", \" \", \"-\", \"--\", \".\", \"n\u002Fa\", \"N\u002FA\", \"na\", \"NA\", \"null\", \"NULL\",\n    \"none\", \"None\", \"unknown\", \"UNKNOWN\", \"tbc\", \"TBC\", \"#N\u002FA\", \"?\",\n}\n\ndef effective_missing(series, numeric_sentinels=()):\n    \"\"\"A boolean mask of values that are missing in substance, not just in form.\"\"\"\n    blank = series.isna()\n\n    if series.dtype == \"object\" or str(series.dtype).startswith(\"string\"):\n        text = series.astype(\"string\").str.strip()\n        blank = blank | text.isin(TEXT_SENTINELS)\n    elif numeric_sentinels:\n        blank = blank | series.isin(list(numeric_sentinels))\n\n    return blank\n\nSENTINELS = {\"units\": (-1, 0), \"revenue\": (-1,)}\n\nfor name in df.columns:\n    mask = effective_missing(df[name], SENTINELS.get(name, ()))\n    print(f\"{name:\u003C12} pandas sees {df[name].isna().sum()}, \"\n          f\"actually missing {int(mask.sum())}\")\n",[14,618,619,629,633,644,700,742,747,751,768,773,783,787,816,831,851,860,880,885,894,899,939,944,959,975,1016],{"__ignoreMap":144},[148,620,621,623,625,627],{"class":150,"line":151},[148,622,184],{"class":183},[148,624,202],{"class":187},[148,626,191],{"class":183},[148,628,207],{"class":187},[148,630,631],{"class":150,"line":197},[148,632,214],{"emptyLinePlaceholder":213},[148,634,635,638,641],{"class":150,"line":210},[148,636,637],{"class":232},"TEXT_SENTINELS",[148,639,640],{"class":183}," =",[148,642,643],{"class":187}," {\n",[148,645,646,649,651,654,656,658,660,663,665,668,670,672,674,677,679,682,684,687,689,692,694,697],{"class":150,"line":217},[148,647,648],{"class":158},"    \"\"",[148,650,236],{"class":187},[148,652,653],{"class":158},"\" \"",[148,655,236],{"class":187},[148,657,292],{"class":158},[148,659,236],{"class":187},[148,661,662],{"class":158},"\"--\"",[148,664,236],{"class":187},[148,666,667],{"class":158},"\".\"",[148,669,236],{"class":187},[148,671,278],{"class":158},[148,673,236],{"class":187},[148,675,676],{"class":158},"\"N\u002FA\"",[148,678,236],{"class":187},[148,680,681],{"class":158},"\"na\"",[148,683,236],{"class":187},[148,685,686],{"class":158},"\"NA\"",[148,688,236],{"class":187},[148,690,691],{"class":158},"\"null\"",[148,693,236],{"class":187},[148,695,696],{"class":158},"\"NULL\"",[148,698,699],{"class":187},",\n",[148,701,702,705,707,710,712,715,717,720,722,725,727,730,732,735,737,740],{"class":150,"line":223},[148,703,704],{"class":158},"    \"none\"",[148,706,236],{"class":187},[148,708,709],{"class":158},"\"None\"",[148,711,236],{"class":187},[148,713,714],{"class":158},"\"unknown\"",[148,716,236],{"class":187},[148,718,719],{"class":158},"\"UNKNOWN\"",[148,721,236],{"class":187},[148,723,724],{"class":158},"\"tbc\"",[148,726,236],{"class":187},[148,728,729],{"class":158},"\"TBC\"",[148,731,236],{"class":187},[148,733,734],{"class":158},"\"#N\u002FA\"",[148,736,236],{"class":187},[148,738,739],{"class":158},"\"?\"",[148,741,699],{"class":187},[148,743,744],{"class":150,"line":265},[148,745,746],{"class":187},"}\n",[148,748,749],{"class":150,"line":302},[148,750,214],{"emptyLinePlaceholder":213},[148,752,753,756,760,763,765],{"class":150,"line":341},[148,754,755],{"class":183},"def",[148,757,759],{"class":758},"s_Opv"," effective_missing",[148,761,762],{"class":187},"(series, numeric_sentinels",[148,764,422],{"class":183},[148,766,767],{"class":187},"()):\n",[148,769,770],{"class":150,"line":374},[148,771,772],{"class":158},"    \"\"\"A boolean mask of values that are missing in substance, not just in form.\"\"\"\n",[148,774,775,778,780],{"class":150,"line":407},[148,776,777],{"class":187},"    blank ",[148,779,422],{"class":183},[148,781,782],{"class":187}," series.isna()\n",[148,784,785],{"class":150,"line":553},[148,786,214],{"emptyLinePlaceholder":213},[148,788,789,792,795,798,801,804,807,810,813],{"class":150,"line":567},[148,790,791],{"class":183},"    if",[148,793,794],{"class":187}," series.dtype ",[148,796,797],{"class":183},"==",[148,799,800],{"class":158}," \"object\"",[148,802,803],{"class":183}," or",[148,805,806],{"class":232}," str",[148,808,809],{"class":187},"(series.dtype).startswith(",[148,811,812],{"class":158},"\"string\"",[148,814,815],{"class":187},"):\n",[148,817,818,821,823,826,828],{"class":150,"line":587},[148,819,820],{"class":187},"        text ",[148,822,422],{"class":183},[148,824,825],{"class":187}," series.astype(",[148,827,812],{"class":158},[148,829,830],{"class":187},").str.strip()\n",[148,832,833,836,838,841,844,847,849],{"class":150,"line":592},[148,834,835],{"class":187},"        blank ",[148,837,422],{"class":183},[148,839,840],{"class":187}," blank ",[148,842,843],{"class":183},"|",[148,845,846],{"class":187}," text.isin(",[148,848,637],{"class":232},[148,850,428],{"class":187},[148,852,854,857],{"class":150,"line":853},15,[148,855,856],{"class":183},"    elif",[148,858,859],{"class":187}," numeric_sentinels:\n",[148,861,863,865,867,869,871,874,877],{"class":150,"line":862},16,[148,864,835],{"class":187},[148,866,422],{"class":183},[148,868,840],{"class":187},[148,870,843],{"class":183},[148,872,873],{"class":187}," series.isin(",[148,875,876],{"class":232},"list",[148,878,879],{"class":187},"(numeric_sentinels))\n",[148,881,883],{"class":150,"line":882},17,[148,884,214],{"emptyLinePlaceholder":213},[148,886,888,891],{"class":150,"line":887},18,[148,889,890],{"class":183},"    return",[148,892,893],{"class":187}," blank\n",[148,895,897],{"class":150,"line":896},19,[148,898,214],{"emptyLinePlaceholder":213},[148,900,902,905,907,910,913,916,918,920,922,924,927,930,932,934,936],{"class":150,"line":901},20,[148,903,904],{"class":232},"SENTINELS",[148,906,640],{"class":183},[148,908,909],{"class":187}," {",[148,911,912],{"class":158},"\"units\"",[148,914,915],{"class":187},": (",[148,917,333],{"class":183},[148,919,336],{"class":232},[148,921,236],{"class":187},[148,923,51],{"class":232},[148,925,926],{"class":187},"), ",[148,928,929],{"class":158},"\"revenue\"",[148,931,915],{"class":187},[148,933,333],{"class":183},[148,935,336],{"class":232},[148,937,938],{"class":187},",)}\n",[148,940,942],{"class":150,"line":941},21,[148,943,214],{"emptyLinePlaceholder":213},[148,945,947,950,953,956],{"class":150,"line":946},22,[148,948,949],{"class":183},"for",[148,951,952],{"class":187}," name ",[148,954,955],{"class":183},"in",[148,957,958],{"class":187}," df.columns:\n",[148,960,962,965,967,970,972],{"class":150,"line":961},23,[148,963,964],{"class":187},"    mask ",[148,966,422],{"class":183},[148,968,969],{"class":187}," effective_missing(df[name], ",[148,971,904],{"class":232},[148,973,974],{"class":187},".get(name, ()))\n",[148,976,978,981,984,987,990,994,997,1000,1003,1006,1008,1011,1013],{"class":150,"line":977},24,[148,979,980],{"class":232},"    print",[148,982,983],{"class":187},"(",[148,985,986],{"class":183},"f",[148,988,989],{"class":158},"\"",[148,991,993],{"class":992},"sSjpA","{",[148,995,996],{"class":187},"name",[148,998,999],{"class":183},":\u003C12",[148,1001,1002],{"class":992},"}",[148,1004,1005],{"class":158}," pandas sees ",[148,1007,993],{"class":992},[148,1009,1010],{"class":187},"df[name].isna().sum()",[148,1012,1002],{"class":992},[148,1014,1015],{"class":158},", \"\n",[148,1017,1019,1022,1025,1027,1030,1033,1035,1037],{"class":150,"line":1018},25,[148,1020,1021],{"class":183},"          f",[148,1023,1024],{"class":158},"\"actually missing ",[148,1026,993],{"class":992},[148,1028,1029],{"class":232},"int",[148,1031,1032],{"class":187},"(mask.sum())",[148,1034,1002],{"class":992},[148,1036,989],{"class":158},[148,1038,428],{"class":187},[10,1040,1041],{},"The gap between the two numbers is the point of the exercise. A column pandas calls 100% complete can be a third empty in substance.",[10,1043,1044,1045,1048],{},"You can also stop pandas guessing on the way in, which matters when a legitimate value collides with a default sentinel — the country code ",[14,1046,1047],{},"NA"," for Namibia is the classic case:",[139,1050,1052],{"className":174,"code":1051,"language":176,"meta":144,"style":144},"import pandas as pd\n\n# Namibia's code survives; only genuinely blank cells become NaN.\ndf = pd.read_excel(\n    \"orders.xlsx\",\n    keep_default_na=False,\n    na_values=[\"\", \" \"],\n)\n",[14,1053,1054,1064,1068,1074,1083,1090,1101,1120],{"__ignoreMap":144},[148,1055,1056,1058,1060,1062],{"class":150,"line":151},[148,1057,184],{"class":183},[148,1059,202],{"class":187},[148,1061,191],{"class":183},[148,1063,207],{"class":187},[148,1065,1066],{"class":150,"line":197},[148,1067,214],{"emptyLinePlaceholder":213},[148,1069,1070],{"class":150,"line":210},[148,1071,1073],{"class":1072},"s-wDw","# Namibia's code survives; only genuinely blank cells become NaN.\n",[148,1075,1076,1078,1080],{"class":150,"line":217},[148,1077,462],{"class":187},[148,1079,422],{"class":183},[148,1081,1082],{"class":187}," pd.read_excel(\n",[148,1084,1085,1088],{"class":150,"line":223},[148,1086,1087],{"class":158},"    \"orders.xlsx\"",[148,1089,699],{"class":187},[148,1091,1092,1095,1097,1099],{"class":150,"line":265},[148,1093,1094],{"class":418},"    keep_default_na",[148,1096,422],{"class":183},[148,1098,425],{"class":232},[148,1100,699],{"class":187},[148,1102,1103,1106,1108,1111,1114,1116,1118],{"class":150,"line":302},[148,1104,1105],{"class":418},"    na_values",[148,1107,422],{"class":183},[148,1109,1110],{"class":187},"[",[148,1112,1113],{"class":158},"\"\"",[148,1115,236],{"class":187},[148,1117,653],{"class":158},[148,1119,262],{"class":187},[148,1121,1122],{"class":150,"line":341},[148,1123,428],{"class":187},[134,1125,1127],{"id":1126},"step-3-separate-spacer-rows-from-real-gaps","Step 3 — Separate spacer rows from real gaps",[10,1129,1130],{},"A blank row in the middle of a sheet is a layout artefact, not a record with missing fields. Counting it as one distorts every column's percentage:",[139,1132,1134],{"className":174,"code":1133,"language":176,"meta":144,"style":144},"import pandas as pd\n\ndef split_blank_rows(df):\n    \"\"\"Return (real rows, blank spacer rows).\"\"\"\n    all_blank = df.isna().all(axis=1)\n    return df.loc[~all_blank].copy(), df.loc[all_blank]\n\ndata, spacers = split_blank_rows(df)\nprint(f\"{len(spacers)} spacer row(s) excluded; {len(data)} real rows\")\n",[14,1135,1136,1146,1150,1160,1165,1184,1197,1201,1211],{"__ignoreMap":144},[148,1137,1138,1140,1142,1144],{"class":150,"line":151},[148,1139,184],{"class":183},[148,1141,202],{"class":187},[148,1143,191],{"class":183},[148,1145,207],{"class":187},[148,1147,1148],{"class":150,"line":197},[148,1149,214],{"emptyLinePlaceholder":213},[148,1151,1152,1154,1157],{"class":150,"line":210},[148,1153,755],{"class":183},[148,1155,1156],{"class":758}," split_blank_rows",[148,1158,1159],{"class":187},"(df):\n",[148,1161,1162],{"class":150,"line":217},[148,1163,1164],{"class":158},"    \"\"\"Return (real rows, blank spacer rows).\"\"\"\n",[148,1166,1167,1170,1172,1175,1178,1180,1182],{"class":150,"line":223},[148,1168,1169],{"class":187},"    all_blank ",[148,1171,422],{"class":183},[148,1173,1174],{"class":187}," df.isna().all(",[148,1176,1177],{"class":418},"axis",[148,1179,422],{"class":183},[148,1181,336],{"class":232},[148,1183,428],{"class":187},[148,1185,1186,1188,1191,1194],{"class":150,"line":265},[148,1187,890],{"class":183},[148,1189,1190],{"class":187}," df.loc[",[148,1192,1193],{"class":183},"~",[148,1195,1196],{"class":187},"all_blank].copy(), df.loc[all_blank]\n",[148,1198,1199],{"class":150,"line":302},[148,1200,214],{"emptyLinePlaceholder":213},[148,1202,1203,1206,1208],{"class":150,"line":341},[148,1204,1205],{"class":187},"data, spacers ",[148,1207,422],{"class":183},[148,1209,1210],{"class":187}," split_blank_rows(df)\n",[148,1212,1213,1215,1217,1219,1221,1223,1225,1228,1230,1233,1235,1237,1240,1242,1245],{"class":150,"line":374},[148,1214,595],{"class":232},[148,1216,983],{"class":187},[148,1218,986],{"class":183},[148,1220,989],{"class":158},[148,1222,993],{"class":992},[148,1224,542],{"class":232},[148,1226,1227],{"class":187},"(spacers)",[148,1229,1002],{"class":992},[148,1231,1232],{"class":158}," spacer row(s) excluded; ",[148,1234,993],{"class":992},[148,1236,542],{"class":232},[148,1238,1239],{"class":187},"(data)",[148,1241,1002],{"class":992},[148,1243,1244],{"class":158}," real rows\"",[148,1246,428],{"class":187},[10,1248,1249],{},"Partly blank rows are the interesting middle case — a record where the key is present but half the fields are absent tells you something different from one where the key itself is gone:",[139,1251,1253],{"className":174,"code":1252,"language":176,"meta":144,"style":144},"import pandas as pd\n\nKEYS = [\"order_id\"]\n\ndef row_completeness(df, keys):\n    \"\"\"Classify rows by how much of them is present.\"\"\"\n    missing_per_row = df.isna().sum(axis=1)\n    key_missing = df[keys].isna().any(axis=1)\n\n    return pd.DataFrame({\n        \"missing_fields\": missing_per_row,\n        \"key_missing\": key_missing,\n        \"class\": pd.cut(\n            missing_per_row \u002F df.shape[1],\n            bins=[-0.01, 0.0, 0.34, 0.67, 1.0],\n            labels=[\"complete\", \"mostly complete\", \"sparse\", \"almost empty\"],\n        ),\n    })\n\nprint(row_completeness(data, KEYS)[\"class\"].value_counts())\n",[14,1254,1255,1265,1269,1285,1289,1299,1304,1322,1340,1344,1350,1358,1366,1374,1389,1425,1454,1459,1464,1468],{"__ignoreMap":144},[148,1256,1257,1259,1261,1263],{"class":150,"line":151},[148,1258,184],{"class":183},[148,1260,202],{"class":187},[148,1262,191],{"class":183},[148,1264,207],{"class":187},[148,1266,1267],{"class":150,"line":197},[148,1268,214],{"emptyLinePlaceholder":213},[148,1270,1271,1274,1276,1279,1282],{"class":150,"line":210},[148,1272,1273],{"class":232},"KEYS",[148,1275,640],{"class":183},[148,1277,1278],{"class":187}," [",[148,1280,1281],{"class":158},"\"order_id\"",[148,1283,1284],{"class":187},"]\n",[148,1286,1287],{"class":150,"line":217},[148,1288,214],{"emptyLinePlaceholder":213},[148,1290,1291,1293,1296],{"class":150,"line":223},[148,1292,755],{"class":183},[148,1294,1295],{"class":758}," row_completeness",[148,1297,1298],{"class":187},"(df, keys):\n",[148,1300,1301],{"class":150,"line":265},[148,1302,1303],{"class":158},"    \"\"\"Classify rows by how much of them is present.\"\"\"\n",[148,1305,1306,1309,1311,1314,1316,1318,1320],{"class":150,"line":302},[148,1307,1308],{"class":187},"    missing_per_row ",[148,1310,422],{"class":183},[148,1312,1313],{"class":187}," df.isna().sum(",[148,1315,1177],{"class":418},[148,1317,422],{"class":183},[148,1319,336],{"class":232},[148,1321,428],{"class":187},[148,1323,1324,1327,1329,1332,1334,1336,1338],{"class":150,"line":341},[148,1325,1326],{"class":187},"    key_missing ",[148,1328,422],{"class":183},[148,1330,1331],{"class":187}," df[keys].isna().any(",[148,1333,1177],{"class":418},[148,1335,422],{"class":183},[148,1337,336],{"class":232},[148,1339,428],{"class":187},[148,1341,1342],{"class":150,"line":374},[148,1343,214],{"emptyLinePlaceholder":213},[148,1345,1346,1348],{"class":150,"line":407},[148,1347,890],{"class":183},[148,1349,523],{"class":187},[148,1351,1352,1355],{"class":150,"line":553},[148,1353,1354],{"class":158},"        \"missing_fields\"",[148,1356,1357],{"class":187},": missing_per_row,\n",[148,1359,1360,1363],{"class":150,"line":567},[148,1361,1362],{"class":158},"        \"key_missing\"",[148,1364,1365],{"class":187},": key_missing,\n",[148,1367,1368,1371],{"class":150,"line":587},[148,1369,1370],{"class":158},"        \"class\"",[148,1372,1373],{"class":187},": pd.cut(\n",[148,1375,1376,1379,1382,1385,1387],{"class":150,"line":592},[148,1377,1378],{"class":187},"            missing_per_row ",[148,1380,1381],{"class":183},"\u002F",[148,1383,1384],{"class":187}," df.shape[",[148,1386,336],{"class":232},[148,1388,262],{"class":187},[148,1390,1391,1394,1396,1398,1400,1403,1405,1408,1410,1413,1415,1418,1420,1423],{"class":150,"line":853},[148,1392,1393],{"class":418},"            bins",[148,1395,422],{"class":183},[148,1397,1110],{"class":187},[148,1399,333],{"class":183},[148,1401,1402],{"class":232},"0.01",[148,1404,236],{"class":187},[148,1406,1407],{"class":232},"0.0",[148,1409,236],{"class":187},[148,1411,1412],{"class":232},"0.34",[148,1414,236],{"class":187},[148,1416,1417],{"class":232},"0.67",[148,1419,236],{"class":187},[148,1421,1422],{"class":232},"1.0",[148,1424,262],{"class":187},[148,1426,1427,1430,1432,1434,1437,1439,1442,1444,1447,1449,1452],{"class":150,"line":862},[148,1428,1429],{"class":418},"            labels",[148,1431,422],{"class":183},[148,1433,1110],{"class":187},[148,1435,1436],{"class":158},"\"complete\"",[148,1438,236],{"class":187},[148,1440,1441],{"class":158},"\"mostly complete\"",[148,1443,236],{"class":187},[148,1445,1446],{"class":158},"\"sparse\"",[148,1448,236],{"class":187},[148,1450,1451],{"class":158},"\"almost empty\"",[148,1453,262],{"class":187},[148,1455,1456],{"class":150,"line":882},[148,1457,1458],{"class":187},"        ),\n",[148,1460,1461],{"class":150,"line":887},[148,1462,1463],{"class":187},"    })\n",[148,1465,1466],{"class":150,"line":896},[148,1467,214],{"emptyLinePlaceholder":213},[148,1469,1470,1472,1475,1477,1480,1483],{"class":150,"line":901},[148,1471,595],{"class":232},[148,1473,1474],{"class":187},"(row_completeness(data, ",[148,1476,1273],{"class":232},[148,1478,1479],{"class":187},")[",[148,1481,1482],{"class":158},"\"class\"",[148,1484,1485],{"class":187},"].value_counts())\n",[10,1487,1488,1489,27],{},"The rows with a missing key are the ones to quarantine rather than fill — a record you cannot identify cannot be reconciled with anything, and the discussion of what to do next belongs in ",[23,1490,1492],{"href":1491},"\u002Fadvanced-data-transformation-and-cleaning\u002Fhandling-missing-data-in-excel-reports\u002Ffill-missing-values-in-excel-with-pandas-fillna\u002F","filling missing values with pandas fillna",[134,1494,1496],{"id":1495},"step-4-look-at-the-pattern-not-just-the-count","Step 4 — Look at the pattern, not just the count",[10,1498,1499],{},"Where the gaps sit matters more than how many there are. Three patterns tell three different stories:",[29,1501,39,1507,39,1510,39,1513,39,1516,39,1522,39,1527,39,1532,39,1564,39,1569,39,1574,39,1584,39,1587,39,1590,39,1601,39,1604],{"viewBox":1502,"role":32,"ariaLabel":1503,"ariaLabelledBy":1504,"xmlns":37,"style":38},"0 0 800 244","Three missing-data patterns: scattered gaps suggesting data-entry lapses, a solid block suggesting a failed batch, and a whole tail suggesting a truncated extract.",[1505,1506],"pat-t","pat-d",[41,1508,1509],{"id":1505},"Three shapes of missingness, three different causes",[45,1511,1512],{"id":1506},"Three grids of rows and columns with missing cells marked. Scattered individual gaps across the grid suggest ordinary data-entry lapses, and filling or excluding them is reasonable. A solid contiguous block in one column over a range of rows suggests one batch or one source failed, which is a pipeline problem rather than a data-quality one. A whole tail of rows missing across every column suggests the extract was truncated, and the missing rows are not missing values at all — they were never delivered.",[49,1514],{"x":51,"y":51,"width":52,"height":1515,"fill":54},"244",[66,1517,1521],{"x":1518,"y":1519,"style":1520},"130","30","font-size:12px;font-weight:700;fill:var(--teal-ink,#0b6157);text-anchor:middle","scattered",[66,1523,1526],{"x":1524,"y":1519,"style":1525},"400","font-size:12px;font-weight:700;fill:var(--gold-ink,#7a4e06);text-anchor:middle","a block",[66,1528,1531],{"x":1529,"y":1519,"style":1530},"670","font-size:12px;font-weight:700;fill:var(--accent-ink,#be185d);text-anchor:middle","a tail",[1533,1534,1535,1536,1535,1544,1535,1551,1535,1554,1535,1558,1535,1561,39],"g",{},"\n    ",[49,1537],{"x":1538,"y":1539,"width":1540,"height":1541,"rx":1542,"fill":54,"stroke":1543,"style":64},"46","44","168","118","8","var(--line,#cdd5e6)",[49,1545],{"x":1546,"y":1547,"width":1548,"height":57,"rx":1549,"fill":1550},"58","56","18","2","#0f9488",[49,1552],{"x":1541,"y":1553,"width":1548,"height":57,"rx":1549,"fill":1550},"76",[49,1555],{"x":1556,"y":1557,"width":1548,"height":57,"rx":1549,"fill":1550},"178","96",[49,1559],{"x":315,"y":1560,"width":1548,"height":57,"rx":1549,"fill":1550},"116",[49,1562],{"x":1563,"y":101,"width":1548,"height":57,"rx":1549,"fill":1550},"148",[66,1565,1568],{"x":1518,"y":1566,"style":1567},"186","font-size:10.5px;fill:var(--text,#172033);text-anchor:middle","data-entry lapses",[66,1570,1573],{"x":1518,"y":1571,"style":1572},"208","font-size:10.5px;fill:var(--muted,#5b6780);text-anchor:middle","fill or exclude row by row",[1533,1575,1535,1576,1535,1579,39],{},[49,1577],{"x":1578,"y":1539,"width":1540,"height":1541,"rx":1542,"fill":54,"stroke":1543,"style":64},"316",[49,1580],{"x":1581,"y":1547,"width":1548,"height":1582,"rx":1549,"fill":1583},"388","94","#b4740a",[66,1585,1586],{"x":1524,"y":1566,"style":1567},"one source or batch failed",[66,1588,1589],{"x":1524,"y":1571,"style":1572},"a pipeline problem, not a data one",[1533,1591,1535,1592,1535,1595,39],{},[49,1593],{"x":1594,"y":1539,"width":1540,"height":1541,"rx":1542,"fill":54,"stroke":1543,"style":64},"586",[49,1596],{"x":1597,"y":1560,"width":1598,"height":1599,"rx":1549,"fill":1600},"598","144","34","#f43f8f",[66,1602,1603],{"x":1529,"y":1566,"style":1567},"the extract was truncated",[66,1605,1606],{"x":1529,"y":1571,"style":1572},"those rows were never delivered",[139,1608,1610],{"className":174,"code":1609,"language":176,"meta":144,"style":144},"import pandas as pd\n\ndef missing_pattern(df, column, bins=20):\n    \"\"\"Where in the file do this column's gaps sit?\"\"\"\n    mask = df[column].isna()\n    if not mask.any():\n        return \"complete\"\n\n    position = pd.cut(pd.Series(range(len(df))), bins=bins, labels=False)\n    by_bin = mask.groupby(position).mean()\n\n    if by_bin.tail(max(1, bins \u002F\u002F 5)).mean() > 0.9:\n        return \"tail — the extract may be truncated\"\n    if (by_bin > 0.9).sum() >= 2 and (by_bin \u003C 0.1).sum() >= 2:\n        return \"block — one batch or source appears to have failed\"\n    return \"scattered — ordinary data-entry gaps\"\n\nfor name in df.columns:\n    print(f\"{name:\u003C12} {missing_pattern(df, name)}\")\n",[14,1611,1612,1622,1626,1643,1648,1657,1667,1675,1679,1716,1726,1730,1765,1772,1811,1818,1825,1829,1839],{"__ignoreMap":144},[148,1613,1614,1616,1618,1620],{"class":150,"line":151},[148,1615,184],{"class":183},[148,1617,202],{"class":187},[148,1619,191],{"class":183},[148,1621,207],{"class":187},[148,1623,1624],{"class":150,"line":197},[148,1625,214],{"emptyLinePlaceholder":213},[148,1627,1628,1630,1633,1636,1638,1641],{"class":150,"line":210},[148,1629,755],{"class":183},[148,1631,1632],{"class":758}," missing_pattern",[148,1634,1635],{"class":187},"(df, column, bins",[148,1637,422],{"class":183},[148,1639,1640],{"class":232},"20",[148,1642,815],{"class":187},[148,1644,1645],{"class":150,"line":217},[148,1646,1647],{"class":158},"    \"\"\"Where in the file do this column's gaps sit?\"\"\"\n",[148,1649,1650,1652,1654],{"class":150,"line":223},[148,1651,964],{"class":187},[148,1653,422],{"class":183},[148,1655,1656],{"class":187}," df[column].isna()\n",[148,1658,1659,1661,1664],{"class":150,"line":265},[148,1660,791],{"class":183},[148,1662,1663],{"class":183}," not",[148,1665,1666],{"class":187}," mask.any():\n",[148,1668,1669,1672],{"class":150,"line":302},[148,1670,1671],{"class":183},"        return",[148,1673,1674],{"class":158}," \"complete\"\n",[148,1676,1677],{"class":150,"line":341},[148,1678,214],{"emptyLinePlaceholder":213},[148,1680,1681,1684,1686,1689,1692,1694,1696,1699,1702,1704,1707,1710,1712,1714],{"class":150,"line":374},[148,1682,1683],{"class":187},"    position ",[148,1685,422],{"class":183},[148,1687,1688],{"class":187}," pd.cut(pd.Series(",[148,1690,1691],{"class":232},"range",[148,1693,983],{"class":187},[148,1695,542],{"class":232},[148,1697,1698],{"class":187},"(df))), ",[148,1700,1701],{"class":418},"bins",[148,1703,422],{"class":183},[148,1705,1706],{"class":187},"bins, ",[148,1708,1709],{"class":418},"labels",[148,1711,422],{"class":183},[148,1713,425],{"class":232},[148,1715,428],{"class":187},[148,1717,1718,1721,1723],{"class":150,"line":407},[148,1719,1720],{"class":187},"    by_bin ",[148,1722,422],{"class":183},[148,1724,1725],{"class":187}," mask.groupby(position).mean()\n",[148,1727,1728],{"class":150,"line":553},[148,1729,214],{"emptyLinePlaceholder":213},[148,1731,1732,1734,1737,1740,1742,1744,1747,1750,1753,1756,1759,1762],{"class":150,"line":567},[148,1733,791],{"class":183},[148,1735,1736],{"class":187}," by_bin.tail(",[148,1738,1739],{"class":232},"max",[148,1741,983],{"class":187},[148,1743,336],{"class":232},[148,1745,1746],{"class":187},", bins ",[148,1748,1749],{"class":183},"\u002F\u002F",[148,1751,1752],{"class":232}," 5",[148,1754,1755],{"class":187},")).mean() ",[148,1757,1758],{"class":183},">",[148,1760,1761],{"class":232}," 0.9",[148,1763,1764],{"class":187},":\n",[148,1766,1767,1769],{"class":150,"line":587},[148,1768,1671],{"class":183},[148,1770,1771],{"class":158}," \"tail — the extract may be truncated\"\n",[148,1773,1774,1776,1779,1781,1783,1786,1789,1792,1795,1797,1800,1803,1805,1807,1809],{"class":150,"line":592},[148,1775,791],{"class":183},[148,1777,1778],{"class":187}," (by_bin ",[148,1780,1758],{"class":183},[148,1782,1761],{"class":232},[148,1784,1785],{"class":187},").sum() ",[148,1787,1788],{"class":183},">=",[148,1790,1791],{"class":232}," 2",[148,1793,1794],{"class":183}," and",[148,1796,1778],{"class":187},[148,1798,1799],{"class":183},"\u003C",[148,1801,1802],{"class":232}," 0.1",[148,1804,1785],{"class":187},[148,1806,1788],{"class":183},[148,1808,1791],{"class":232},[148,1810,1764],{"class":187},[148,1812,1813,1815],{"class":150,"line":853},[148,1814,1671],{"class":183},[148,1816,1817],{"class":158}," \"block — one batch or source appears to have failed\"\n",[148,1819,1820,1822],{"class":150,"line":862},[148,1821,890],{"class":183},[148,1823,1824],{"class":158}," \"scattered — ordinary data-entry gaps\"\n",[148,1826,1827],{"class":150,"line":882},[148,1828,214],{"emptyLinePlaceholder":213},[148,1830,1831,1833,1835,1837],{"class":150,"line":887},[148,1832,949],{"class":183},[148,1834,952],{"class":187},[148,1836,955],{"class":183},[148,1838,958],{"class":187},[148,1840,1841,1843,1845,1847,1849,1851,1853,1855,1857,1859,1862,1864,1866],{"class":150,"line":896},[148,1842,980],{"class":232},[148,1844,983],{"class":187},[148,1846,986],{"class":183},[148,1848,989],{"class":158},[148,1850,993],{"class":992},[148,1852,996],{"class":187},[148,1854,999],{"class":183},[148,1856,1002],{"class":992},[148,1858,909],{"class":992},[148,1860,1861],{"class":187},"missing_pattern(df, name)",[148,1863,1002],{"class":992},[148,1865,989],{"class":158},[148,1867,428],{"class":187},[10,1869,1870],{},"Distinguishing these is what turns \"the revenue column is 30% empty\" into an actionable statement. A tail means asking the sender to re-run the extract; scattered gaps mean deciding a fill strategy.",[134,1872,1874],{"id":1873},"step-5-write-the-report","Step 5 — Write the report",[29,1876,39,1882,39,1885,39,1888,39,1891,39,1895,39,1900,39,1905,39,1911,39,1914,39,1920,39,1926,39,1929,39,1932,39,1935,39,1938,39,1942,39,1947,39,1950,39,1954,39,1957,39,1960,39,1965,39,1968,39,1971,39,1975,39,1978,39,1980,39,1983,39,1986],{"viewBox":1877,"role":32,"ariaLabel":1878,"ariaLabelledBy":1879,"xmlns":37,"style":38},"0 0 800 232","Per-column expectations instead of one threshold: a key column held to one hundred per cent, a revenue column to ninety, and a notes column with no expectation at all.",[1880,1881],"exp-t","exp-d",[41,1883,1884],{"id":1880},"One threshold for the file flags the wrong columns",[45,1886,1887],{"id":1881},"Four columns with their actual completeness shown as bars against a per-column expectation marker. The order id column is ninety-eight per cent complete against a hundred per cent expectation and is therefore a genuine failure. Revenue at ninety-four per cent against ninety passes. Region at ninety-six against ninety-five passes. The notes column at four per cent has no expectation and passes trivially. A single file-wide threshold of ninety would have flagged notes and missed order id — exactly backwards.",[49,1889],{"x":51,"y":51,"width":52,"height":1890,"fill":54},"232",[66,1892,1894],{"x":1524,"y":58,"style":1893},"font-size:12px;font-weight:700;fill:var(--muted,#5b6780);text-anchor:middle","completeness against a per-column expectation",[49,1896],{"x":1897,"y":1539,"width":1898,"height":58,"rx":1899,"fill":120,"stroke":121},"16","112","5",[66,1901,1904],{"x":1902,"y":1903,"style":131},"72","62","order_id",[49,1906],{"x":1598,"y":1538,"width":1907,"height":1908,"rx":1909,"fill":1910},"560","22","4","#f0f2f5",[49,1912],{"x":1598,"y":1538,"width":1913,"height":1908,"rx":1909,"fill":1600},"549",[150,1915],{"x1":1916,"y1":1917,"x2":1916,"y2":1902,"stroke":1918,"style":1919},"704","42","var(--text,#172033)","stroke-width:2.5px",[66,1921,1925],{"x":1922,"y":1923,"style":1924},"716","63","font-size:10.5px;font-weight:700;fill:var(--accent-ink,#be185d)","FAIL",[49,1927],{"x":1897,"y":1928,"width":1898,"height":58,"rx":1899,"fill":120,"stroke":121},"84",[66,1930,1931],{"x":1902,"y":79,"style":131},"revenue",[49,1933],{"x":1598,"y":1934,"width":1907,"height":1908,"rx":1909,"fill":1910},"86",[49,1936],{"x":1598,"y":1934,"width":1937,"height":1908,"rx":1909,"fill":1550},"526",[150,1939],{"x1":1940,"y1":1941,"x2":1940,"y2":1898,"stroke":1918,"style":1919},"648","82",[66,1943,1946],{"x":1922,"y":1944,"style":1945},"103","font-size:10.5px;font-weight:700;fill:var(--teal-ink,#0b6157)","OK",[49,1948],{"x":1897,"y":1949,"width":1898,"height":58,"rx":1899,"fill":120,"stroke":121},"124",[66,1951,1953],{"x":1902,"y":1952,"style":131},"142","region",[49,1955],{"x":1598,"y":1956,"width":1907,"height":1908,"rx":1909,"fill":1910},"126",[49,1958],{"x":1598,"y":1956,"width":1959,"height":1908,"rx":1909,"fill":1550},"538",[150,1961],{"x1":1962,"y1":1963,"x2":1962,"y2":1964,"stroke":1918,"style":1919},"676","122","152",[66,1966,1946],{"x":1922,"y":1967,"style":1945},"143",[49,1969],{"x":1897,"y":1970,"width":1898,"height":58,"rx":1899,"fill":120,"stroke":121},"164",[66,1972,1974],{"x":1902,"y":1973,"style":131},"182","note",[49,1976],{"x":1598,"y":1977,"width":1907,"height":1908,"rx":1909,"fill":1910},"166",[49,1979],{"x":1598,"y":1977,"width":1908,"height":1908,"rx":1909,"fill":1550},[150,1981],{"x1":1598,"y1":106,"x2":1598,"y2":1982,"stroke":1918,"style":1919},"192",[66,1984,1946],{"x":1922,"y":1985,"style":1945},"183",[66,1987,1990],{"x":1524,"y":1988,"style":1989},"218","font-size:11px;fill:var(--muted,#5b6780);text-anchor:middle","a single 90% threshold would flag \"note\" and miss \"order_id\" — exactly backwards",[10,1992,1993],{},"Produce something a non-programmer can open, with the gaps highlighted:",[139,1995,1997],{"className":174,"code":1996,"language":176,"meta":144,"style":144},"import pandas as pd\n\ndef completeness_report(path, dest, sentinels=None, expectations=None):\n    \"\"\"Audit an Excel file and write a formatted completeness report.\"\"\"\n    df = pd.read_excel(path)\n    data, spacers = split_blank_rows(df)\n    sentinels = sentinels or {}\n    expectations = expectations or {}\n\n    rows = []\n    for name in data.columns:\n        mask = effective_missing(data[name], sentinels.get(name, ()))\n        missing = int(mask.sum())\n        complete = round((1 - missing \u002F len(data)) * 100, 1) if len(data) else 0.0\n        required = expectations.get(name, 0.0)\n        rows.append({\n            \"column\": name,\n            \"rows\": len(data),\n            \"missing\": missing,\n            \"complete_pct\": complete,\n            \"required_pct\": required,\n            \"status\": \"OK\" if complete >= required else \"BELOW EXPECTATION\",\n            \"pattern\": missing_pattern(data, name),\n        })\n\n    report = pd.DataFrame(rows).sort_values(\"complete_pct\")\n\n    with pd.ExcelWriter(dest, engine=\"xlsxwriter\") as writer:\n        report.to_excel(writer, sheet_name=\"Completeness\", index=False)\n        book, sheet = writer.book, writer.sheets[\"Completeness\"]\n\n        header = book.add_format({\"bold\": True, \"bg_color\": \"#EEF2FF\",\n                                  \"border\": 1})\n        for position, name in enumerate(report.columns):\n            sheet.write(0, position, name, header)\n\n        bad = book.add_format({\"bg_color\": \"#FEE8F2\", \"font_color\": \"#BE185D\"})\n        sheet.conditional_format(\n            1, 5, len(report), 5,\n            {\"type\": \"text\", \"criteria\": \"containing\",\n             \"value\": \"BELOW\", \"format\": bad},\n        )\n        sheet.set_column(\"A:A\", 20)\n        sheet.set_column(\"B:F\", 14)\n        sheet.set_column(\"G:G\", 42)\n        sheet.freeze_panes(1, 0)\n\n    return report\n\ncompleteness_report(\n    \"orders.xlsx\", \"completeness.xlsx\",\n    sentinels={\"units\": (-1, 0)},\n    expectations={\"order_id\": 100.0, \"region\": 95.0, \"revenue\": 90.0},\n)\n",[14,1998,1999,2009,2013,2036,2041,2051,2060,2076,2090,2094,2104,2116,2126,2139,2193,2207,2212,2220,2232,2239,2247,2255,2283,2291,2296,2300,2316,2321,2345,2369,2384,2389,2420,2433,2450,2461,2466,2495,2501,2522,2548,2567,2573,2588,2602,2616,2630,2635,2643,2648,2654,2666,2691,2730],{"__ignoreMap":144},[148,2000,2001,2003,2005,2007],{"class":150,"line":151},[148,2002,184],{"class":183},[148,2004,202],{"class":187},[148,2006,191],{"class":183},[148,2008,207],{"class":187},[148,2010,2011],{"class":150,"line":197},[148,2012,214],{"emptyLinePlaceholder":213},[148,2014,2015,2017,2020,2023,2025,2027,2030,2032,2034],{"class":150,"line":210},[148,2016,755],{"class":183},[148,2018,2019],{"class":758}," completeness_report",[148,2021,2022],{"class":187},"(path, dest, sentinels",[148,2024,422],{"class":183},[148,2026,249],{"class":232},[148,2028,2029],{"class":187},", expectations",[148,2031,422],{"class":183},[148,2033,249],{"class":232},[148,2035,815],{"class":187},[148,2037,2038],{"class":150,"line":217},[148,2039,2040],{"class":158},"    \"\"\"Audit an Excel file and write a formatted completeness report.\"\"\"\n",[148,2042,2043,2046,2048],{"class":150,"line":223},[148,2044,2045],{"class":187},"    df ",[148,2047,422],{"class":183},[148,2049,2050],{"class":187}," pd.read_excel(path)\n",[148,2052,2053,2056,2058],{"class":150,"line":265},[148,2054,2055],{"class":187},"    data, spacers ",[148,2057,422],{"class":183},[148,2059,1210],{"class":187},[148,2061,2062,2065,2067,2070,2073],{"class":150,"line":302},[148,2063,2064],{"class":187},"    sentinels ",[148,2066,422],{"class":183},[148,2068,2069],{"class":187}," sentinels ",[148,2071,2072],{"class":183},"or",[148,2074,2075],{"class":187}," {}\n",[148,2077,2078,2081,2083,2086,2088],{"class":150,"line":341},[148,2079,2080],{"class":187},"    expectations ",[148,2082,422],{"class":183},[148,2084,2085],{"class":187}," expectations ",[148,2087,2072],{"class":183},[148,2089,2075],{"class":187},[148,2091,2092],{"class":150,"line":374},[148,2093,214],{"emptyLinePlaceholder":213},[148,2095,2096,2099,2101],{"class":150,"line":407},[148,2097,2098],{"class":187},"    rows ",[148,2100,422],{"class":183},[148,2102,2103],{"class":187}," []\n",[148,2105,2106,2109,2111,2113],{"class":150,"line":553},[148,2107,2108],{"class":183},"    for",[148,2110,952],{"class":187},[148,2112,955],{"class":183},[148,2114,2115],{"class":187}," data.columns:\n",[148,2117,2118,2121,2123],{"class":150,"line":567},[148,2119,2120],{"class":187},"        mask ",[148,2122,422],{"class":183},[148,2124,2125],{"class":187}," effective_missing(data[name], sentinels.get(name, ()))\n",[148,2127,2128,2131,2133,2136],{"class":150,"line":587},[148,2129,2130],{"class":187},"        missing ",[148,2132,422],{"class":183},[148,2134,2135],{"class":232}," int",[148,2137,2138],{"class":187},"(mask.sum())\n",[148,2140,2141,2144,2146,2149,2152,2154,2156,2159,2161,2164,2167,2169,2172,2174,2176,2179,2182,2184,2187,2190],{"class":150,"line":592},[148,2142,2143],{"class":187},"        complete ",[148,2145,422],{"class":183},[148,2147,2148],{"class":232}," round",[148,2150,2151],{"class":187},"((",[148,2153,336],{"class":232},[148,2155,500],{"class":183},[148,2157,2158],{"class":187}," missing ",[148,2160,1381],{"class":183},[148,2162,2163],{"class":232}," len",[148,2165,2166],{"class":187},"(data)) ",[148,2168,506],{"class":183},[148,2170,2171],{"class":232}," 100",[148,2173,236],{"class":187},[148,2175,336],{"class":232},[148,2177,2178],{"class":187},") ",[148,2180,2181],{"class":183},"if",[148,2183,2163],{"class":232},[148,2185,2186],{"class":187},"(data) ",[148,2188,2189],{"class":183},"else",[148,2191,2192],{"class":232}," 0.0\n",[148,2194,2195,2198,2200,2203,2205],{"class":150,"line":853},[148,2196,2197],{"class":187},"        required ",[148,2199,422],{"class":183},[148,2201,2202],{"class":187}," expectations.get(name, ",[148,2204,1407],{"class":232},[148,2206,428],{"class":187},[148,2208,2209],{"class":150,"line":862},[148,2210,2211],{"class":187},"        rows.append({\n",[148,2213,2214,2217],{"class":150,"line":882},[148,2215,2216],{"class":158},"            \"column\"",[148,2218,2219],{"class":187},": name,\n",[148,2221,2222,2225,2227,2229],{"class":150,"line":887},[148,2223,2224],{"class":158},"            \"rows\"",[148,2226,539],{"class":187},[148,2228,542],{"class":232},[148,2230,2231],{"class":187},"(data),\n",[148,2233,2234,2237],{"class":150,"line":896},[148,2235,2236],{"class":158},"            \"missing\"",[148,2238,531],{"class":187},[148,2240,2241,2244],{"class":150,"line":901},[148,2242,2243],{"class":158},"            \"complete_pct\"",[148,2245,2246],{"class":187},": complete,\n",[148,2248,2249,2252],{"class":150,"line":941},[148,2250,2251],{"class":158},"            \"required_pct\"",[148,2253,2254],{"class":187},": required,\n",[148,2256,2257,2260,2262,2265,2268,2271,2273,2276,2278,2281],{"class":150,"line":946},[148,2258,2259],{"class":158},"            \"status\"",[148,2261,539],{"class":187},[148,2263,2264],{"class":158},"\"OK\"",[148,2266,2267],{"class":183}," if",[148,2269,2270],{"class":187}," complete ",[148,2272,1788],{"class":183},[148,2274,2275],{"class":187}," required ",[148,2277,2189],{"class":183},[148,2279,2280],{"class":158}," \"BELOW EXPECTATION\"",[148,2282,699],{"class":187},[148,2284,2285,2288],{"class":150,"line":961},[148,2286,2287],{"class":158},"            \"pattern\"",[148,2289,2290],{"class":187},": missing_pattern(data, name),\n",[148,2292,2293],{"class":150,"line":977},[148,2294,2295],{"class":187},"        })\n",[148,2297,2298],{"class":150,"line":1018},[148,2299,214],{"emptyLinePlaceholder":213},[148,2301,2303,2306,2308,2311,2314],{"class":150,"line":2302},26,[148,2304,2305],{"class":187},"    report ",[148,2307,422],{"class":183},[148,2309,2310],{"class":187}," pd.DataFrame(rows).sort_values(",[148,2312,2313],{"class":158},"\"complete_pct\"",[148,2315,428],{"class":187},[148,2317,2319],{"class":150,"line":2318},27,[148,2320,214],{"emptyLinePlaceholder":213},[148,2322,2324,2327,2330,2333,2335,2338,2340,2342],{"class":150,"line":2323},28,[148,2325,2326],{"class":183},"    with",[148,2328,2329],{"class":187}," pd.ExcelWriter(dest, ",[148,2331,2332],{"class":418},"engine",[148,2334,422],{"class":183},[148,2336,2337],{"class":158},"\"xlsxwriter\"",[148,2339,2178],{"class":187},[148,2341,191],{"class":183},[148,2343,2344],{"class":187}," writer:\n",[148,2346,2348,2351,2354,2356,2359,2361,2363,2365,2367],{"class":150,"line":2347},29,[148,2349,2350],{"class":187},"        report.to_excel(writer, ",[148,2352,2353],{"class":418},"sheet_name",[148,2355,422],{"class":183},[148,2357,2358],{"class":158},"\"Completeness\"",[148,2360,236],{"class":187},[148,2362,419],{"class":418},[148,2364,422],{"class":183},[148,2366,425],{"class":232},[148,2368,428],{"class":187},[148,2370,2372,2375,2377,2380,2382],{"class":150,"line":2371},30,[148,2373,2374],{"class":187},"        book, sheet ",[148,2376,422],{"class":183},[148,2378,2379],{"class":187}," writer.book, writer.sheets[",[148,2381,2358],{"class":158},[148,2383,1284],{"class":187},[148,2385,2387],{"class":150,"line":2386},31,[148,2388,214],{"emptyLinePlaceholder":213},[148,2390,2392,2395,2397,2400,2403,2405,2408,2410,2413,2415,2418],{"class":150,"line":2391},32,[148,2393,2394],{"class":187},"        header ",[148,2396,422],{"class":183},[148,2398,2399],{"class":187}," book.add_format({",[148,2401,2402],{"class":158},"\"bold\"",[148,2404,539],{"class":187},[148,2406,2407],{"class":232},"True",[148,2409,236],{"class":187},[148,2411,2412],{"class":158},"\"bg_color\"",[148,2414,539],{"class":187},[148,2416,2417],{"class":158},"\"#EEF2FF\"",[148,2419,699],{"class":187},[148,2421,2423,2426,2428,2430],{"class":150,"line":2422},33,[148,2424,2425],{"class":158},"                                  \"border\"",[148,2427,539],{"class":187},[148,2429,336],{"class":232},[148,2431,2432],{"class":187},"})\n",[148,2434,2436,2439,2442,2444,2447],{"class":150,"line":2435},34,[148,2437,2438],{"class":183},"        for",[148,2440,2441],{"class":187}," position, name ",[148,2443,955],{"class":183},[148,2445,2446],{"class":232}," enumerate",[148,2448,2449],{"class":187},"(report.columns):\n",[148,2451,2453,2456,2458],{"class":150,"line":2452},35,[148,2454,2455],{"class":187},"            sheet.write(",[148,2457,51],{"class":232},[148,2459,2460],{"class":187},", position, name, header)\n",[148,2462,2464],{"class":150,"line":2463},36,[148,2465,214],{"emptyLinePlaceholder":213},[148,2467,2469,2472,2474,2476,2478,2480,2483,2485,2488,2490,2493],{"class":150,"line":2468},37,[148,2470,2471],{"class":187},"        bad ",[148,2473,422],{"class":183},[148,2475,2399],{"class":187},[148,2477,2412],{"class":158},[148,2479,539],{"class":187},[148,2481,2482],{"class":158},"\"#FEE8F2\"",[148,2484,236],{"class":187},[148,2486,2487],{"class":158},"\"font_color\"",[148,2489,539],{"class":187},[148,2491,2492],{"class":158},"\"#BE185D\"",[148,2494,2432],{"class":187},[148,2496,2498],{"class":150,"line":2497},38,[148,2499,2500],{"class":187},"        sheet.conditional_format(\n",[148,2502,2504,2507,2509,2511,2513,2515,2518,2520],{"class":150,"line":2503},39,[148,2505,2506],{"class":232},"            1",[148,2508,236],{"class":187},[148,2510,1899],{"class":232},[148,2512,236],{"class":187},[148,2514,542],{"class":232},[148,2516,2517],{"class":187},"(report), ",[148,2519,1899],{"class":232},[148,2521,699],{"class":187},[148,2523,2525,2528,2531,2533,2536,2538,2541,2543,2546],{"class":150,"line":2524},40,[148,2526,2527],{"class":187},"            {",[148,2529,2530],{"class":158},"\"type\"",[148,2532,539],{"class":187},[148,2534,2535],{"class":158},"\"text\"",[148,2537,236],{"class":187},[148,2539,2540],{"class":158},"\"criteria\"",[148,2542,539],{"class":187},[148,2544,2545],{"class":158},"\"containing\"",[148,2547,699],{"class":187},[148,2549,2551,2554,2556,2559,2561,2564],{"class":150,"line":2550},41,[148,2552,2553],{"class":158},"             \"value\"",[148,2555,539],{"class":187},[148,2557,2558],{"class":158},"\"BELOW\"",[148,2560,236],{"class":187},[148,2562,2563],{"class":158},"\"format\"",[148,2565,2566],{"class":187},": bad},\n",[148,2568,2570],{"class":150,"line":2569},42,[148,2571,2572],{"class":187},"        )\n",[148,2574,2576,2579,2582,2584,2586],{"class":150,"line":2575},43,[148,2577,2578],{"class":187},"        sheet.set_column(",[148,2580,2581],{"class":158},"\"A:A\"",[148,2583,236],{"class":187},[148,2585,1640],{"class":232},[148,2587,428],{"class":187},[148,2589,2591,2593,2596,2598,2600],{"class":150,"line":2590},44,[148,2592,2578],{"class":187},[148,2594,2595],{"class":158},"\"B:F\"",[148,2597,236],{"class":187},[148,2599,57],{"class":232},[148,2601,428],{"class":187},[148,2603,2605,2607,2610,2612,2614],{"class":150,"line":2604},45,[148,2606,2578],{"class":187},[148,2608,2609],{"class":158},"\"G:G\"",[148,2611,236],{"class":187},[148,2613,1917],{"class":232},[148,2615,428],{"class":187},[148,2617,2619,2622,2624,2626,2628],{"class":150,"line":2618},46,[148,2620,2621],{"class":187},"        sheet.freeze_panes(",[148,2623,336],{"class":232},[148,2625,236],{"class":187},[148,2627,51],{"class":232},[148,2629,428],{"class":187},[148,2631,2633],{"class":150,"line":2632},47,[148,2634,214],{"emptyLinePlaceholder":213},[148,2636,2638,2640],{"class":150,"line":2637},48,[148,2639,890],{"class":183},[148,2641,2642],{"class":187}," report\n",[148,2644,2646],{"class":150,"line":2645},49,[148,2647,214],{"emptyLinePlaceholder":213},[148,2649,2651],{"class":150,"line":2650},50,[148,2652,2653],{"class":187},"completeness_report(\n",[148,2655,2657,2659,2661,2664],{"class":150,"line":2656},51,[148,2658,1087],{"class":158},[148,2660,236],{"class":187},[148,2662,2663],{"class":158},"\"completeness.xlsx\"",[148,2665,699],{"class":187},[148,2667,2669,2672,2674,2676,2678,2680,2682,2684,2686,2688],{"class":150,"line":2668},52,[148,2670,2671],{"class":418},"    sentinels",[148,2673,422],{"class":183},[148,2675,993],{"class":187},[148,2677,912],{"class":158},[148,2679,915],{"class":187},[148,2681,333],{"class":183},[148,2683,336],{"class":232},[148,2685,236],{"class":187},[148,2687,51],{"class":232},[148,2689,2690],{"class":187},")},\n",[148,2692,2694,2697,2699,2701,2703,2705,2708,2710,2713,2715,2718,2720,2722,2724,2727],{"class":150,"line":2693},53,[148,2695,2696],{"class":418},"    expectations",[148,2698,422],{"class":183},[148,2700,993],{"class":187},[148,2702,1281],{"class":158},[148,2704,539],{"class":187},[148,2706,2707],{"class":232},"100.0",[148,2709,236],{"class":187},[148,2711,2712],{"class":158},"\"region\"",[148,2714,539],{"class":187},[148,2716,2717],{"class":232},"95.0",[148,2719,236],{"class":187},[148,2721,929],{"class":158},[148,2723,539],{"class":187},[148,2725,2726],{"class":232},"90.0",[148,2728,2729],{"class":187},"},\n",[148,2731,2733],{"class":150,"line":2732},54,[148,2734,428],{"class":187},[10,2736,2737,2738,27],{},"Per-column expectations are what makes the report useful. A blanket threshold flags the commentary column that is meant to be empty and misses the key column that is 2% short — which is the one that matters. The highlighting technique generalises, as shown in ",[23,2739,2741],{"href":2740},"\u002Fadvanced-data-transformation-and-cleaning\u002Fvalidating-excel-data-with-python\u002Fhighlight-invalid-cells-in-excel-with-python\u002F","highlighting invalid cells in Excel with Python",[134,2743,2745],{"id":2744},"common-pitfalls-and-fixes","Common pitfalls and fixes",[2747,2748,2749,2765],"table",{},[2750,2751,2752],"thead",{},[2753,2754,2755,2759,2762],"tr",{},[2756,2757,2758],"th",{},"Symptom",[2756,2760,2761],{},"Cause",[2756,2763,2764],{},"Fix",[2766,2767,2768,2780,2800,2811,2825,2836,2847],"tbody",{},[2753,2769,2770,2774,2777],{},[2771,2772,2773],"td",{},"Column reports 100% complete but is not",[2771,2775,2776],{},"Text sentinels read as data",[2771,2778,2779],{},"Check against an explicit sentinel list.",[2753,2781,2782,2788,2794],{},[2771,2783,2784,2785,2787],{},"Country code ",[14,2786,1047],{}," became blank",[2771,2789,2790,2791],{},"pandas default ",[14,2792,2793],{},"na_values",[2771,2795,2796,2799],{},[14,2797,2798],{},"keep_default_na=False"," plus your own list.",[2753,2801,2802,2805,2808],{},[2771,2803,2804],{},"Every column looks 10% empty",[2771,2806,2807],{},"Spacer rows counted as records",[2771,2809,2810],{},"Exclude all-blank rows first.",[2753,2812,2813,2816,2822],{},[2771,2814,2815],{},"Averages look wrong",[2771,2817,2818,2819,2821],{},"Numeric sentinel such as ",[14,2820,20],{}," included",[2771,2823,2824],{},"Treat sentinels as missing before aggregating.",[2753,2826,2827,2830,2833],{},[2771,2828,2829],{},"Report flags a notes column",[2771,2831,2832],{},"Blanket threshold",[2771,2834,2835],{},"Set expectations per column.",[2753,2837,2838,2841,2844],{},[2771,2839,2840],{},"Missing values reappear next run",[2771,2842,2843],{},"Fixed the symptom, not the source",[2771,2845,2846],{},"Report the pattern; escalate a block or tail.",[2753,2848,2849,2852,2855],{},[2771,2850,2851],{},"A whole tail of rows is empty",[2771,2853,2854],{},"Extract truncated upstream",[2771,2856,2857],{},"Ask for a re-run; do not fill.",[134,2859,2861],{"id":2860},"performance-and-scale-notes","Performance and scale notes",[10,2863,2864,2867],{},[14,2865,2866],{},"isna()"," builds a boolean frame the same shape as the data, so auditing a large workbook briefly doubles its memory footprint. Two adjustments keep that manageable.",[10,2869,2870,2874,2875,2878],{},[2871,2872,2873],"strong",{},"Aggregate per column rather than materialising the whole mask."," ",[14,2876,2877],{},"df.isna().sum()"," builds the full frame; a loop over columns builds one column at a time:",[139,2880,2882],{"className":174,"code":2881,"language":176,"meta":144,"style":144},"missing = {name: int(df[name].isna().sum()) for name in df.columns}\n",[14,2883,2884],{"__ignoreMap":144},[148,2885,2886,2888,2890,2893,2895,2898,2900,2902,2904],{"class":150,"line":151},[148,2887,480],{"class":187},[148,2889,422],{"class":183},[148,2891,2892],{"class":187}," {name: ",[148,2894,1029],{"class":232},[148,2896,2897],{"class":187},"(df[name].isna().sum()) ",[148,2899,949],{"class":183},[148,2901,952],{"class":187},[148,2903,955],{"class":183},[148,2905,2906],{"class":187}," df.columns}\n",[10,2908,2909,2912],{},[2871,2910,2911],{},"Audit a sample for a first pass."," Completeness percentages stabilise quickly, so a sample of a hundred thousand rows gives a reliable picture of a ten-million-row file at a fraction of the cost. Follow up with a full pass only on the columns the sample flagged.",[10,2914,2915],{},"For files too large to hold at all, accumulate the counts chunk by chunk — the counts are additive, so partial results combine cleanly:",[139,2917,2919],{"className":174,"code":2918,"language":176,"meta":144,"style":144},"import pandas as pd\n\ntotals, rows = None, 0\nfor chunk in pd.read_csv(\"huge_export.csv\", chunksize=200_000):\n    part = chunk.isna().sum()\n    totals = part if totals is None else totals.add(part, fill_value=0)\n    rows += len(chunk)\n\nprint((1 - totals \u002F rows).mul(100).round(1).sort_values().to_string())\n",[14,2920,2921,2931,2935,2950,2977,2987,3022,3034,3038],{"__ignoreMap":144},[148,2922,2923,2925,2927,2929],{"class":150,"line":151},[148,2924,184],{"class":183},[148,2926,202],{"class":187},[148,2928,191],{"class":183},[148,2930,207],{"class":187},[148,2932,2933],{"class":150,"line":197},[148,2934,214],{"emptyLinePlaceholder":213},[148,2936,2937,2940,2942,2945,2947],{"class":150,"line":210},[148,2938,2939],{"class":187},"totals, rows ",[148,2941,422],{"class":183},[148,2943,2944],{"class":232}," None",[148,2946,236],{"class":187},[148,2948,2949],{"class":232},"0\n",[148,2951,2952,2954,2957,2959,2962,2965,2967,2970,2972,2975],{"class":150,"line":217},[148,2953,949],{"class":183},[148,2955,2956],{"class":187}," chunk ",[148,2958,955],{"class":183},[148,2960,2961],{"class":187}," pd.read_csv(",[148,2963,2964],{"class":158},"\"huge_export.csv\"",[148,2966,236],{"class":187},[148,2968,2969],{"class":418},"chunksize",[148,2971,422],{"class":183},[148,2973,2974],{"class":232},"200_000",[148,2976,815],{"class":187},[148,2978,2979,2982,2984],{"class":150,"line":223},[148,2980,2981],{"class":187},"    part ",[148,2983,422],{"class":183},[148,2985,2986],{"class":187}," chunk.isna().sum()\n",[148,2988,2989,2992,2994,2997,2999,3002,3005,3007,3010,3013,3016,3018,3020],{"class":150,"line":265},[148,2990,2991],{"class":187},"    totals ",[148,2993,422],{"class":183},[148,2995,2996],{"class":187}," part ",[148,2998,2181],{"class":183},[148,3000,3001],{"class":187}," totals ",[148,3003,3004],{"class":183},"is",[148,3006,2944],{"class":232},[148,3008,3009],{"class":183}," else",[148,3011,3012],{"class":187}," totals.add(part, ",[148,3014,3015],{"class":418},"fill_value",[148,3017,422],{"class":183},[148,3019,51],{"class":232},[148,3021,428],{"class":187},[148,3023,3024,3026,3029,3031],{"class":150,"line":302},[148,3025,2098],{"class":187},[148,3027,3028],{"class":183},"+=",[148,3030,2163],{"class":232},[148,3032,3033],{"class":187},"(chunk)\n",[148,3035,3036],{"class":150,"line":341},[148,3037,214],{"emptyLinePlaceholder":213},[148,3039,3040,3042,3044,3046,3048,3050,3052,3055,3058,3061,3063],{"class":150,"line":374},[148,3041,595],{"class":232},[148,3043,2151],{"class":187},[148,3045,336],{"class":232},[148,3047,500],{"class":183},[148,3049,3001],{"class":187},[148,3051,1381],{"class":183},[148,3053,3054],{"class":187}," rows).mul(",[148,3056,3057],{"class":232},"100",[148,3059,3060],{"class":187},").round(",[148,3062,336],{"class":232},[148,3064,3065],{"class":187},").sort_values().to_string())\n",[10,3067,3068,3069,3073],{},"The same chunked shape works for Excel via the approach in ",[23,3070,3072],{"href":3071},"\u002Fadvanced-data-transformation-and-cleaning\u002Fworking-with-large-excel-files-in-python\u002Fread-large-excel-file-in-chunks-with-pandas\u002F","reading large Excel files in chunks",", and running the audit at ingest — before anything downstream depends on the data — is where it costs least and catches most.",[134,3075,3077],{"id":3076},"conclusion","Conclusion",[10,3079,3080,3081,3083,3084,3086],{},"A completeness audit is ",[14,3082,437],{}," plus everything ",[14,3085,437],{}," cannot see. List the text sentinels the file actually uses, name the numeric ones per column, and exclude whole-blank spacer rows so the percentages mean something. Then look at where the gaps sit: scattered gaps are a data-quality question, a solid block is a failed source, and a missing tail means the extract was truncated and no amount of filling will help. Set expectations per column rather than one threshold for the file, and write the result somewhere a human will read it before anybody fills anything.",[134,3088,3090],{"id":3089},"frequently-asked-questions","Frequently asked questions",[10,3092,3093,3096,3097,236,3100,236,3102,3105,3106,236,3108,236,3111,3114,3115,3118,3119,3122,3123,3125],{},[2871,3094,3095],{},"Which values does pandas treat as missing by default?","\nEmpty cells, ",[14,3098,3099],{},"NaN",[14,3101,249],{},[14,3103,3104],{},"NaT"," and a short list of strings including ",[14,3107,1047],{},[14,3109,3110],{},"N\u002FA",[14,3112,3113],{},"null"," and ",[14,3116,3117],{},"nan",". Anything else — a dash, the word ",[14,3120,3121],{},"unknown",", a single space, the number ",[14,3124,20],{}," used as a sentinel — is read as ordinary data.",[10,3127,3128,3131,3132,3134,3135,3137,3138,3140,3141,3143],{},[2871,3129,3130],{},"How do I stop pandas treating a real value as missing?","\nPass ",[14,3133,2798],{}," and supply your own ",[14,3136,2793],{}," list. That matters for a genuine product code like ",[14,3139,1047],{}," or a country code like ",[14,3142,1047],{}," for Namibia, which the defaults would otherwise blank out.",[10,3145,3146,3149],{},[2871,3147,3148],{},"What is a good completeness threshold?","\nIt depends on the column, not the file. A key column should be one hundred per cent complete, a commentary column can be almost entirely empty. Set the expectation per column and check against it.",[10,3151,3152,3155],{},[2871,3153,3154],{},"Should I report missing values or just fill them?","\nReport first, always. Filling before understanding turns a broken upstream extract into a plausible-looking report, and the fill choice depends on why the values are missing in the first place.",[10,3157,3158,3161,3162,3165],{},[2871,3159,3160],{},"How do I find rows that are entirely blank?","\nUse ",[14,3163,3164],{},"df.isna().all(axis=1)"," to select them. They usually come from spacer rows in the source sheet rather than from real records, so counting them separately keeps the per-column figures honest.",[134,3167,3169],{"id":3168},"related","Related",[3171,3172,3173,3180,3186,3193,3200],"ul",{},[3174,3175,3176,3177,3179],"li",{},"Up to the parent: ",[23,3178,26],{"href":25}," — what to do once you know what is missing.",[3174,3181,3182,3185],{},[23,3183,3184],{"href":1491},"Fill Missing Values in Excel with pandas fillna"," — the fill strategies this audit informs.",[3174,3187,3188,3192],{},[23,3189,3191],{"href":3190},"\u002Fadvanced-data-transformation-and-cleaning\u002Fhandling-missing-data-in-excel-reports\u002Finterpolate-missing-numeric-values-in-excel-data\u002F","Interpolate Missing Numeric Values in Excel Data"," — filling gaps in an ordered series.",[3174,3194,3195,3199],{},[23,3196,3198],{"href":3197},"\u002Fadvanced-data-transformation-and-cleaning\u002Fcleaning-excel-data-with-pandas\u002Fremove-blank-rows-from-excel-with-pandas\u002F","Remove Blank Rows from Excel with pandas"," — dealing with the spacer rows.",[3174,3201,3202,3205],{},[23,3203,3204],{"href":2740},"Highlight Invalid Cells in Excel with Python"," — showing the gaps in the workbook itself.",[3207,3208,3209],"style",{},"html pre.shiki code .sMTad, html code.shiki .sMTad{--shiki-default:#6F42C1;--shiki-dark:#FFB757}html pre.shiki code .srMev, html code.shiki .srMev{--shiki-default:#032F62;--shiki-dark:#ADDCFF}html .default .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html.dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html pre.shiki code .s-kum, html code.shiki .s-kum{--shiki-default:#D73A49;--shiki-dark:#FF9492}html pre.shiki code .skGVy, html code.shiki .skGVy{--shiki-default:#24292E;--shiki-dark:#F0F3F6}html pre.shiki code .sP0c6, html code.shiki .sP0c6{--shiki-default:#005CC5;--shiki-dark:#91CBFF}html pre.shiki code .sa561, html code.shiki .sa561{--shiki-default:#E36209;--shiki-dark:#FFB757}html pre.shiki code .s_Opv, html code.shiki .s_Opv{--shiki-default:#6F42C1;--shiki-dark:#DBB7FF}html pre.shiki code .sSjpA, html code.shiki .sSjpA{--shiki-default:#005CC5;--shiki-dark:#FF9492}html pre.shiki code .s-wDw, html code.shiki .s-wDw{--shiki-default:#6A737D;--shiki-dark:#BDC4CC}",{"title":144,"searchDepth":197,"depth":197,"links":3211},[3212,3213,3214,3215,3216,3217,3218,3219,3220,3221,3222],{"id":136,"depth":197,"text":137},{"id":431,"depth":197,"text":432},{"id":609,"depth":197,"text":610},{"id":1126,"depth":197,"text":1127},{"id":1495,"depth":197,"text":1496},{"id":1873,"depth":197,"text":1874},{"id":2744,"depth":197,"text":2745},{"id":2860,"depth":197,"text":2861},{"id":3076,"depth":197,"text":3077},{"id":3089,"depth":197,"text":3090},{"id":3168,"depth":197,"text":3169},"2026-08-15","Audit an Excel file for gaps before you use it — per-column null counts, disguised blanks like n\u002Fa and dashes, missing-value patterns, and a formatted report of what is absent.","md",[3227,3229,3231,3233,3235],{"q":3095,"a":3228},"Empty cells, NaN, None, NaT and a short list of strings including NA, N\u002FA, null and nan. Anything else — a dash, the word unknown, a single space, the number minus one used as a sentinel — is read as ordinary data.",{"q":3130,"a":3230},"Pass keep_default_na=False and supply your own na_values list. That matters for a genuine product code like NA or a country code like NA for Namibia, which the defaults would otherwise blank out.",{"q":3148,"a":3232},"It depends on the column, not the file. A key column should be one hundred per cent complete, a commentary column can be almost entirely empty. Set the expectation per column and check against it.",{"q":3154,"a":3234},"Report first, always. Filling before understanding turns a broken upstream extract into a plausible-looking report, and the fill choice depends on why the values are missing in the first place.",{"q":3160,"a":3236},"Use df.isna().all(axis=1) to select them. They usually come from spacer rows in the source sheet rather than from real records, so counting them separately keeps the per-column figures honest.",{},"\u002Fadvanced-data-transformation-and-cleaning\u002Fhandling-missing-data-in-excel-reports\u002Ffind-and-report-missing-values-in-an-excel-file",{"title":3240,"description":3241},"Find Missing Values in an Excel File with pandas","Profile an Excel workbook's completeness in Python: isna counts, sentinel values pandas does not recognise, whole-blank rows, per-column thresholds and an exportable report.","find-and-report-missing-values-in-an-excel-file","advanced-data-transformation-and-cleaning\u002Fhandling-missing-data-in-excel-reports\u002Ffind-and-report-missing-values-in-an-excel-file\u002Findex","how-to","vF6ZyKrlBSeB5A_buMqzzQMaD-makuSyTL269Dll2Sc",[3247,3251],{"title":3248,"path":3249,"stem":3250,"children":-1},"Fill Missing Values in Excel with Pandas fillna","\u002Fadvanced-data-transformation-and-cleaning\u002Fhandling-missing-data-in-excel-reports\u002Ffill-missing-values-in-excel-with-pandas-fillna","advanced-data-transformation-and-cleaning\u002Fhandling-missing-data-in-excel-reports\u002Ffill-missing-values-in-excel-with-pandas-fillna\u002Findex",{"title":3191,"path":3252,"stem":3253,"children":-1},"\u002Fadvanced-data-transformation-and-cleaning\u002Fhandling-missing-data-in-excel-reports\u002Finterpolate-missing-numeric-values-in-excel-data","advanced-data-transformation-and-cleaning\u002Fhandling-missing-data-in-excel-reports\u002Finterpolate-missing-numeric-values-in-excel-data\u002Findex",1786800027129]