[{"data":1,"prerenderedAt":3154},["ShallowReactive",2],{"doc:\u002Fadvanced-data-transformation-and-cleaning\u002Fmerging-and-joining-excel-dataframes\u002Ffind-rows-in-one-excel-file-missing-from-another":3,"surround:\u002Fadvanced-data-transformation-and-cleaning\u002Fmerging-and-joining-excel-dataframes\u002Ffind-rows-in-one-excel-file-missing-from-another":3146},{"id":4,"title":5,"body":6,"dateModified":3120,"datePublished":3120,"description":3121,"extension":3122,"faq":3123,"meta":3137,"navigation":257,"path":3138,"seo":3139,"slug":3142,"stem":3143,"type":3144,"__hash__":3145},"docs\u002Fadvanced-data-transformation-and-cleaning\u002Fmerging-and-joining-excel-dataframes\u002Ffind-rows-in-one-excel-file-missing-from-another\u002Findex.md","Find Rows in One Excel File Missing from Another",{"type":7,"value":8,"toc":3107},"minimark",[9,19,189,194,226,229,477,491,495,498,712,723,833,840,844,945,948,1117,1120,1124,1127,1361,1373,1377,1384,1466,1943,1946,1950,1953,2501,2508,2512,2623,2627,2633,2640,2709,2715,2724,2803,2806,2991,3003,3007,3013,3017,3035,3041,3050,3056,3062,3066,3103],[10,11,12,13,18],"p",{},"Two files that should agree, and a number that does not. The question is always the same: which rows are in one and not the other, and which are in both but different? pandas answers it with a single merge — as long as the keys match, which on real spreadsheet data they frequently do not. This guide covers the reconciliation itself, the invisible key mismatches that make it lie, and the exceptions workbook that turns the result into something a colleague can act on. It extends ",[14,15,17],"a",{"href":16},"\u002Fadvanced-data-transformation-and-cleaning\u002Fmerging-and-joining-excel-dataframes\u002F","Merging and Joining Excel DataFrames",".",[20,21,30,31,30,35,30,39,30,46,30,56,30,63,30,68,30,73,30,79,30,84,30,88,30,101,30,106,30,112,30,118,30,123,30,134,30,146,30,155,30,161,30,165,30,168,30,173,30,177,30,180,30,185],"svg",{"viewBox":22,"role":23,"ariaLabel":24,"ariaLabelledBy":25,"xmlns":28,"style":29},"0 0 800 254","img","An outer merge with an indicator splitting two files into three groups: rows only in the first, rows in both, and rows only in the second.",[26,27],"rec-t","rec-d","http:\u002F\u002Fwww.w3.org\u002F2000\u002Fsvg","width:100%;max-width:800px;height:auto;display:block;margin:1.5rem auto;font-family:Inter,ui-sans-serif,system-ui,sans-serif","\n  ",[32,33,34],"title",{"id":26},"One outer merge answers both directions at once",[36,37,38],"desc",{"id":27},"Two source files overlap partially. An outer merge with the indicator flag labels every resulting row as left_only, both, or right_only. Left_only rows exist in the first file and are missing from the second. Right_only rows are the reverse. Rows marked both are present in each and can be compared value by value to find changes. A left join would answer only one of the three questions.",[40,41],"rect",{"x":42,"y":42,"width":43,"height":44,"fill":45},"0","800","254","#ffffff",[40,47],{"x":48,"y":49,"width":50,"height":51,"rx":52,"fill":53,"stroke":54,"style":55},"14","72","150","86","13","#ebebfd","var(--brand,#5b5cf0)","stroke-width:2px",[57,58,62],"text",{"x":59,"y":60,"style":61},"89","56","font-size:11px;font-weight:700;fill:var(--muted,#5b6780);text-anchor:middle","file A",[57,64,67],{"x":59,"y":65,"style":66},"108","font-size:12px;font-weight:700;fill:var(--brand-strong,#4338ca);text-anchor:middle","system export",[57,69,72],{"x":59,"y":70,"style":71},"130","font-size:10.5px;fill:var(--muted,#5b6780);text-anchor:middle","1,482 rows",[40,74],{"x":48,"y":75,"width":50,"height":76,"rx":52,"fill":77,"stroke":78,"style":55},"176","62","#d9f4f1","var(--teal,#0f9488)",[57,80,83],{"x":59,"y":81,"style":82},"204","font-size:12px;font-weight:700;fill:var(--teal-ink,#0b6157);text-anchor:middle","file B",[57,85,87],{"x":59,"y":86,"style":71},"224","1,470 rows",[89,90,93,94,93,98,30],"g",{"stroke":91,"style":55,"fill":92},"var(--line,#cdd5e6)","none","\n    ",[95,96],"path",{"d":97},"M164 115 H 200 V 140 H 232",[95,99],{"d":100},"M164 207 H 200 V 140",[102,103],"polygon",{"points":104,"fill":105},"240,140 228,134 228,146","#5b5cf0",[40,107],{"x":108,"y":109,"width":75,"height":49,"rx":52,"fill":110,"stroke":111,"style":55},"248","104","#fdefd8","var(--gold,#b4740a)",[57,113,117],{"x":114,"y":115,"style":116},"336","132","font-size:11.5px;font-weight:700;fill:var(--gold-ink,#7a4e06);text-anchor:middle","outer merge",[57,119,122],{"x":114,"y":120,"style":121},"154","font-size:10.5px;fill:var(--text,#172033);text-anchor:middle","indicator=True",[89,124,93,125,93,128,93,131,30],{"stroke":111,"style":55,"fill":92},[95,126],{"d":127},"M424 140 H 458 V 52 H 494",[95,129],{"d":130},"M424 140 H 462",[95,132],{"d":133},"M424 140 H 458 V 218 H 494",[89,135,93,137,93,140,93,143,30],{"fill":136},"#b4740a",[102,138],{"points":139},"502,52 490,46 490,58",[102,141],{"points":142},"502,140 490,134 490,146",[102,144],{"points":145},"502,218 490,212 490,224",[40,147],{"x":148,"y":149,"width":150,"height":151,"rx":152,"fill":153,"stroke":154,"style":55},"510","28","274","52","11","#fee8f2","var(--accent,#f43f8f)",[57,156,160],{"x":157,"y":158,"style":159},"647","50","font-size:11.5px;font-weight:700;fill:var(--accent-ink,#be185d);text-anchor:middle","left_only — missing from B",[57,162,164],{"x":157,"y":163,"style":71},"69","18 rows",[40,166],{"x":148,"y":167,"width":150,"height":151,"rx":152,"fill":77,"stroke":78,"style":55},"116",[57,169,172],{"x":157,"y":170,"style":171},"138","font-size:11.5px;font-weight:700;fill:var(--teal-ink,#0b6157);text-anchor:middle","both — compare the values",[57,174,176],{"x":157,"y":175,"style":71},"157","1,464 rows · some may differ",[40,178],{"x":148,"y":179,"width":150,"height":151,"rx":152,"fill":53,"stroke":54,"style":55},"194",[57,181,184],{"x":157,"y":182,"style":183},"216","font-size:11.5px;font-weight:700;fill:var(--brand-strong,#4338ca);text-anchor:middle","right_only — missing from A",[57,186,188],{"x":157,"y":187,"style":71},"235","6 rows",[190,191,193],"h2",{"id":192},"prerequisites","Prerequisites",[195,196,201],"pre",{"className":197,"code":198,"language":199,"meta":200,"style":200},"language-bash shiki shiki-themes github-light github-dark-high-contrast","pip install pandas openpyxl xlsxwriter\n","bash","",[202,203,204],"code",{"__ignoreMap":200},[205,206,209,213,217,220,223],"span",{"class":207,"line":208},"line",1,[205,210,212],{"class":211},"sMTad","pip",[205,214,216],{"class":215},"srMev"," install",[205,218,219],{"class":215}," pandas",[205,221,222],{"class":215}," openpyxl",[205,224,225],{"class":215}," xlsxwriter\n",[10,227,228],{},"Two files that nearly agree:",[195,230,234],{"className":231,"code":232,"language":233,"meta":200,"style":200},"language-python shiki shiki-themes github-light github-dark-high-contrast","import pandas as pd\n\npd.DataFrame({\n    \"order_id\": [\"A-1001\", \"A-1002\", \"A-1003\", \"A-1004\"],\n    \"region\": [\"North\", \"South\", \"West\", \"North\"],\n    \"revenue\": [5150.00, 4268.50, 3511.25, 2980.10],\n}).to_excel(\"system.xlsx\", index=False)\n\npd.DataFrame({\n    \"order_id\": [\"A-1001\", \"A-1002 \", \"A-1004\", \"A-1005\"],\n    \"region\": [\"North\", \"South\", \"North\", \"East\"],\n    \"revenue\": [5150.00, 4268.50, 2980.10, 1820.00],\n}).to_excel(\"ledger.xlsx\", index=False)\n","python",[202,235,236,252,259,265,296,323,352,376,381,386,411,435,459],{"__ignoreMap":200},[205,237,238,242,246,249],{"class":207,"line":208},[205,239,241],{"class":240},"s-kum","import",[205,243,245],{"class":244},"skGVy"," pandas ",[205,247,248],{"class":240},"as",[205,250,251],{"class":244}," pd\n",[205,253,255],{"class":207,"line":254},2,[205,256,258],{"emptyLinePlaceholder":257},true,"\n",[205,260,262],{"class":207,"line":261},3,[205,263,264],{"class":244},"pd.DataFrame({\n",[205,266,268,271,274,277,280,283,285,288,290,293],{"class":207,"line":267},4,[205,269,270],{"class":215},"    \"order_id\"",[205,272,273],{"class":244},": [",[205,275,276],{"class":215},"\"A-1001\"",[205,278,279],{"class":244},", ",[205,281,282],{"class":215},"\"A-1002\"",[205,284,279],{"class":244},[205,286,287],{"class":215},"\"A-1003\"",[205,289,279],{"class":244},[205,291,292],{"class":215},"\"A-1004\"",[205,294,295],{"class":244},"],\n",[205,297,299,302,304,307,309,312,314,317,319,321],{"class":207,"line":298},5,[205,300,301],{"class":215},"    \"region\"",[205,303,273],{"class":244},[205,305,306],{"class":215},"\"North\"",[205,308,279],{"class":244},[205,310,311],{"class":215},"\"South\"",[205,313,279],{"class":244},[205,315,316],{"class":215},"\"West\"",[205,318,279],{"class":244},[205,320,306],{"class":215},[205,322,295],{"class":244},[205,324,326,329,331,335,337,340,342,345,347,350],{"class":207,"line":325},6,[205,327,328],{"class":215},"    \"revenue\"",[205,330,273],{"class":244},[205,332,334],{"class":333},"sP0c6","5150.00",[205,336,279],{"class":244},[205,338,339],{"class":333},"4268.50",[205,341,279],{"class":244},[205,343,344],{"class":333},"3511.25",[205,346,279],{"class":244},[205,348,349],{"class":333},"2980.10",[205,351,295],{"class":244},[205,353,355,358,361,363,367,370,373],{"class":207,"line":354},7,[205,356,357],{"class":244},"}).to_excel(",[205,359,360],{"class":215},"\"system.xlsx\"",[205,362,279],{"class":244},[205,364,366],{"class":365},"sa561","index",[205,368,369],{"class":240},"=",[205,371,372],{"class":333},"False",[205,374,375],{"class":244},")\n",[205,377,379],{"class":207,"line":378},8,[205,380,258],{"emptyLinePlaceholder":257},[205,382,384],{"class":207,"line":383},9,[205,385,264],{"class":244},[205,387,389,391,393,395,397,400,402,404,406,409],{"class":207,"line":388},10,[205,390,270],{"class":215},[205,392,273],{"class":244},[205,394,276],{"class":215},[205,396,279],{"class":244},[205,398,399],{"class":215},"\"A-1002 \"",[205,401,279],{"class":244},[205,403,292],{"class":215},[205,405,279],{"class":244},[205,407,408],{"class":215},"\"A-1005\"",[205,410,295],{"class":244},[205,412,414,416,418,420,422,424,426,428,430,433],{"class":207,"line":413},11,[205,415,301],{"class":215},[205,417,273],{"class":244},[205,419,306],{"class":215},[205,421,279],{"class":244},[205,423,311],{"class":215},[205,425,279],{"class":244},[205,427,306],{"class":215},[205,429,279],{"class":244},[205,431,432],{"class":215},"\"East\"",[205,434,295],{"class":244},[205,436,438,440,442,444,446,448,450,452,454,457],{"class":207,"line":437},12,[205,439,328],{"class":215},[205,441,273],{"class":244},[205,443,334],{"class":333},[205,445,279],{"class":244},[205,447,339],{"class":333},[205,449,279],{"class":244},[205,451,349],{"class":333},[205,453,279],{"class":244},[205,455,456],{"class":333},"1820.00",[205,458,295],{"class":244},[205,460,462,464,467,469,471,473,475],{"class":207,"line":461},13,[205,463,357],{"class":244},[205,465,466],{"class":215},"\"ledger.xlsx\"",[205,468,279],{"class":244},[205,470,366],{"class":365},[205,472,369],{"class":240},[205,474,372],{"class":333},[205,476,375],{"class":244},[10,478,479,482,483,486,487,490],{},[202,480,481],{},"A-1003"," is genuinely missing from the ledger, ",[202,484,485],{},"A-1005"," is genuinely extra — and ",[202,488,489],{},"A-1002 "," has a trailing space that will make it look missing from both directions unless the keys are normalised.",[190,492,494],{"id":493},"step-1-normalise-the-keys-first","Step 1 — Normalise the keys first",[10,496,497],{},"Skip this and the reconciliation reports differences that are not real:",[195,499,501],{"className":231,"code":500,"language":233,"meta":200,"style":200},"import re\nimport pandas as pd\n\ndef normalise_key(series):\n    \"\"\"A comparison key that survives whitespace, case and formatting drift.\"\"\"\n    text = series.astype(\"string\")\n    text = text.str.replace(\" \", \" \", regex=False)     # non-breaking space\n    text = text.str.replace(r\"\\s+\", \" \", regex=True).str.strip()\n    return text.str.casefold()\n\nsystem = pd.read_excel(\"system.xlsx\")\nledger = pd.read_excel(\"ledger.xlsx\")\n\nsystem[\"key\"] = normalise_key(system[\"order_id\"])\nledger[\"key\"] = normalise_key(ledger[\"order_id\"])\n",[202,502,503,510,520,524,536,541,556,589,627,635,639,653,666,670,693],{"__ignoreMap":200},[205,504,505,507],{"class":207,"line":208},[205,506,241],{"class":240},[205,508,509],{"class":244}," re\n",[205,511,512,514,516,518],{"class":207,"line":254},[205,513,241],{"class":240},[205,515,245],{"class":244},[205,517,248],{"class":240},[205,519,251],{"class":244},[205,521,522],{"class":207,"line":261},[205,523,258],{"emptyLinePlaceholder":257},[205,525,526,529,533],{"class":207,"line":267},[205,527,528],{"class":240},"def",[205,530,532],{"class":531},"s_Opv"," normalise_key",[205,534,535],{"class":244},"(series):\n",[205,537,538],{"class":207,"line":298},[205,539,540],{"class":215},"    \"\"\"A comparison key that survives whitespace, case and formatting drift.\"\"\"\n",[205,542,543,546,548,551,554],{"class":207,"line":325},[205,544,545],{"class":244},"    text ",[205,547,369],{"class":240},[205,549,550],{"class":244}," series.astype(",[205,552,553],{"class":215},"\"string\"",[205,555,375],{"class":244},[205,557,558,560,562,565,568,570,573,575,578,580,582,585],{"class":207,"line":354},[205,559,545],{"class":244},[205,561,369],{"class":240},[205,563,564],{"class":244}," text.str.replace(",[205,566,567],{"class":215},"\" \"",[205,569,279],{"class":244},[205,571,572],{"class":215},"\" \"",[205,574,279],{"class":244},[205,576,577],{"class":365},"regex",[205,579,369],{"class":240},[205,581,372],{"class":333},[205,583,584],{"class":244},")     ",[205,586,588],{"class":587},"s-wDw","# non-breaking space\n",[205,590,591,593,595,597,600,603,606,609,611,613,615,617,619,621,624],{"class":207,"line":378},[205,592,545],{"class":244},[205,594,369],{"class":240},[205,596,564],{"class":244},[205,598,599],{"class":240},"r",[205,601,602],{"class":215},"\"",[205,604,605],{"class":333},"\\s",[205,607,608],{"class":240},"+",[205,610,602],{"class":215},[205,612,279],{"class":244},[205,614,572],{"class":215},[205,616,279],{"class":244},[205,618,577],{"class":365},[205,620,369],{"class":240},[205,622,623],{"class":333},"True",[205,625,626],{"class":244},").str.strip()\n",[205,628,629,632],{"class":207,"line":383},[205,630,631],{"class":240},"    return",[205,633,634],{"class":244}," text.str.casefold()\n",[205,636,637],{"class":207,"line":388},[205,638,258],{"emptyLinePlaceholder":257},[205,640,641,644,646,649,651],{"class":207,"line":413},[205,642,643],{"class":244},"system ",[205,645,369],{"class":240},[205,647,648],{"class":244}," pd.read_excel(",[205,650,360],{"class":215},[205,652,375],{"class":244},[205,654,655,658,660,662,664],{"class":207,"line":437},[205,656,657],{"class":244},"ledger ",[205,659,369],{"class":240},[205,661,648],{"class":244},[205,663,466],{"class":215},[205,665,375],{"class":244},[205,667,668],{"class":207,"line":461},[205,669,258],{"emptyLinePlaceholder":257},[205,671,673,676,679,682,684,687,690],{"class":207,"line":672},14,[205,674,675],{"class":244},"system[",[205,677,678],{"class":215},"\"key\"",[205,680,681],{"class":244},"] ",[205,683,369],{"class":240},[205,685,686],{"class":244}," normalise_key(system[",[205,688,689],{"class":215},"\"order_id\"",[205,691,692],{"class":244},"])\n",[205,694,696,699,701,703,705,708,710],{"class":207,"line":695},15,[205,697,698],{"class":244},"ledger[",[205,700,678],{"class":215},[205,702,681],{"class":244},[205,704,369],{"class":240},[205,706,707],{"class":244}," normalise_key(ledger[",[205,709,689],{"class":215},[205,711,692],{"class":244},[10,713,714,715,718,719,722],{},"Numeric keys read as text on one side and numbers on the other are the other frequent culprit — ",[202,716,717],{},"1001"," and ",[202,720,721],{},"1001.0"," will not match:",[195,724,726],{"className":231,"code":725,"language":233,"meta":200,"style":200},"def normalise_numeric_key(series):\n    \"\"\"Canonicalise an identifier that may arrive as text or as a number.\"\"\"\n    numbers = pd.to_numeric(series, errors=\"coerce\")\n    as_text = series.astype(\"string\").str.strip()\n    # Where it parsed as a whole number, use the integer form.\n    return numbers.map(\n        lambda v: str(int(v)) if pd.notna(v) and float(v).is_integer() else None\n    ).fillna(as_text)\n",[202,727,728,737,742,762,775,780,787,828],{"__ignoreMap":200},[205,729,730,732,735],{"class":207,"line":208},[205,731,528],{"class":240},[205,733,734],{"class":531}," normalise_numeric_key",[205,736,535],{"class":244},[205,738,739],{"class":207,"line":254},[205,740,741],{"class":215},"    \"\"\"Canonicalise an identifier that may arrive as text or as a number.\"\"\"\n",[205,743,744,747,749,752,755,757,760],{"class":207,"line":261},[205,745,746],{"class":244},"    numbers ",[205,748,369],{"class":240},[205,750,751],{"class":244}," pd.to_numeric(series, ",[205,753,754],{"class":365},"errors",[205,756,369],{"class":240},[205,758,759],{"class":215},"\"coerce\"",[205,761,375],{"class":244},[205,763,764,767,769,771,773],{"class":207,"line":267},[205,765,766],{"class":244},"    as_text ",[205,768,369],{"class":240},[205,770,550],{"class":244},[205,772,553],{"class":215},[205,774,626],{"class":244},[205,776,777],{"class":207,"line":298},[205,778,779],{"class":587},"    # Where it parsed as a whole number, use the integer form.\n",[205,781,782,784],{"class":207,"line":325},[205,783,631],{"class":240},[205,785,786],{"class":244}," numbers.map(\n",[205,788,789,792,795,798,801,804,807,810,813,816,819,822,825],{"class":207,"line":354},[205,790,791],{"class":240},"        lambda",[205,793,794],{"class":244}," v: ",[205,796,797],{"class":333},"str",[205,799,800],{"class":244},"(",[205,802,803],{"class":333},"int",[205,805,806],{"class":244},"(v)) ",[205,808,809],{"class":240},"if",[205,811,812],{"class":244}," pd.notna(v) ",[205,814,815],{"class":240},"and",[205,817,818],{"class":333}," float",[205,820,821],{"class":244},"(v).is_integer() ",[205,823,824],{"class":240},"else",[205,826,827],{"class":333}," None\n",[205,829,830],{"class":207,"line":378},[205,831,832],{"class":244},"    ).fillna(as_text)\n",[10,834,835,836,18],{},"The wider treatment of invisible text differences is in ",[14,837,839],{"href":838},"\u002Fadvanced-data-transformation-and-cleaning\u002Fcleaning-excel-data-with-pandas\u002Fstrip-whitespace-and-normalise-text-columns-with-pandas\u002F","stripping whitespace and normalising text columns",[190,841,843],{"id":842},"step-2-check-the-keys-are-unique","Step 2 — Check the keys are unique",[20,845,30,851,30,854,30,857,30,860,30,866,30,871,30,876,30,879,30,883,30,887,30,891,30,895,30,900,30,903,30,907,30,911,30,915,30,919,30,922,30,927,30,933,30,938,30,941],{"viewBox":846,"role":23,"ariaLabel":847,"ariaLabelledBy":848,"xmlns":28,"style":29},"0 0 800 228","A duplicated key multiplies rows: two matching rows on one side and three on the other produce six rows in the merge, making every count meaningless.",[849,850],"dupkey-t","dupkey-d",[32,852,853],{"id":849},"What a duplicated key does to a merge",[36,855,856],{"id":850},"The key A-1002 appears twice in the left file and three times in the right. A merge pairs every left occurrence with every right occurrence, producing six rows from five. The counts that follow are therefore meaningless, and worse, any revenue total computed from the merged frame is inflated. Checking for duplicated keys before merging turns this into an immediate error rather than a wrong number.",[40,858],{"x":42,"y":42,"width":43,"height":859,"fill":45},"228",[57,861,865],{"x":862,"y":863,"style":864},"120","30","font-size:11.5px;font-weight:700;fill:var(--muted,#5b6780);text-anchor:middle","left: 2 rows",[40,867],{"x":863,"y":868,"width":869,"height":863,"rx":870,"fill":53,"stroke":54,"style":55},"44","180","5",[57,872,875],{"x":862,"y":873,"style":874},"64","font-size:10.5px;font-weight:700;fill:var(--brand-strong,#4338ca);text-anchor:middle","A-1002 · 100",[40,877],{"x":863,"y":878,"width":869,"height":863,"rx":870,"fill":53,"stroke":54,"style":55},"80",[57,880,882],{"x":862,"y":881,"style":874},"100","A-1002 · 200",[57,884,886],{"x":862,"y":885,"style":864},"146","right: 3 rows",[40,888],{"x":863,"y":889,"width":869,"height":890,"rx":870,"fill":77,"stroke":78,"style":55},"158","26",[57,892,894],{"x":862,"y":75,"style":893},"font-size:10.5px;font-weight:700;fill:var(--teal-ink,#0b6157);text-anchor:middle","A-1002 × 3",[207,896],{"x1":897,"y1":898,"x2":899,"y2":898,"stroke":54,"style":55},"222","112","266",[102,901],{"points":902,"fill":105},"274,112 262,106 262,118",[40,904],{"x":905,"y":51,"width":906,"height":151,"rx":152,"fill":110,"stroke":111,"style":55},"282","164",[57,908,910],{"x":909,"y":65,"style":116},"364","merge on key",[57,912,914],{"x":909,"y":913,"style":71},"127","every pair matched",[207,916],{"x1":917,"y1":898,"x2":918,"y2":898,"stroke":111,"style":55},"446","490",[102,920],{"points":921,"fill":136},"498,112 486,106 486,118",[40,923],{"x":924,"y":868,"width":925,"height":926,"rx":52,"fill":153,"stroke":154,"style":55},"506","278","136",[57,928,932],{"x":929,"y":930,"style":931},"645","76","font-size:13px;font-weight:700;fill:var(--accent-ink,#be185d);text-anchor:middle","6 rows out of 5",[57,934,937],{"x":929,"y":935,"style":936},"106","font-size:11px;fill:var(--text,#172033);text-anchor:middle","2 × 3 = 6 combinations",[57,939,940],{"x":929,"y":115,"style":936},"every count is now wrong",[57,942,944],{"x":929,"y":889,"style":943},"font-size:11px;font-weight:700;fill:var(--accent-ink,#be185d);text-anchor:middle","and revenue totals are inflated",[10,946,947],{},"A merge on a duplicated key multiplies rows, and the resulting counts are meaningless:",[195,949,951],{"className":231,"code":950,"language":233,"meta":200,"style":200},"import pandas as pd\n\ndef check_unique(df, key, label):\n    duplicated = df[key].duplicated(keep=False)\n    if duplicated.any():\n        counts = df.loc[duplicated, key].value_counts()\n        raise ValueError(\n            f\"{label}: {int(duplicated.sum())} rows share a key. \"\n            f\"Worst offenders:\\n{counts.head().to_string()}\"\n        )\n    return True\n\ncheck_unique(system, \"key\", \"system.xlsx\")\ncheck_unique(ledger, \"key\", \"ledger.xlsx\")\n",[202,952,953,963,967,977,996,1004,1014,1025,1057,1075,1080,1087,1091,1104],{"__ignoreMap":200},[205,954,955,957,959,961],{"class":207,"line":208},[205,956,241],{"class":240},[205,958,245],{"class":244},[205,960,248],{"class":240},[205,962,251],{"class":244},[205,964,965],{"class":207,"line":254},[205,966,258],{"emptyLinePlaceholder":257},[205,968,969,971,974],{"class":207,"line":261},[205,970,528],{"class":240},[205,972,973],{"class":531}," check_unique",[205,975,976],{"class":244},"(df, key, label):\n",[205,978,979,982,984,987,990,992,994],{"class":207,"line":267},[205,980,981],{"class":244},"    duplicated ",[205,983,369],{"class":240},[205,985,986],{"class":244}," df[key].duplicated(",[205,988,989],{"class":365},"keep",[205,991,369],{"class":240},[205,993,372],{"class":333},[205,995,375],{"class":244},[205,997,998,1001],{"class":207,"line":298},[205,999,1000],{"class":240},"    if",[205,1002,1003],{"class":244}," duplicated.any():\n",[205,1005,1006,1009,1011],{"class":207,"line":325},[205,1007,1008],{"class":244},"        counts ",[205,1010,369],{"class":240},[205,1012,1013],{"class":244}," df.loc[duplicated, key].value_counts()\n",[205,1015,1016,1019,1022],{"class":207,"line":354},[205,1017,1018],{"class":240},"        raise",[205,1020,1021],{"class":333}," ValueError",[205,1023,1024],{"class":244},"(\n",[205,1026,1027,1030,1032,1036,1039,1042,1045,1047,1049,1052,1054],{"class":207,"line":378},[205,1028,1029],{"class":240},"            f",[205,1031,602],{"class":215},[205,1033,1035],{"class":1034},"sSjpA","{",[205,1037,1038],{"class":244},"label",[205,1040,1041],{"class":1034},"}",[205,1043,1044],{"class":215},": ",[205,1046,1035],{"class":1034},[205,1048,803],{"class":333},[205,1050,1051],{"class":244},"(duplicated.sum())",[205,1053,1041],{"class":1034},[205,1055,1056],{"class":215}," rows share a key. \"\n",[205,1058,1059,1061,1064,1067,1070,1072],{"class":207,"line":383},[205,1060,1029],{"class":240},[205,1062,1063],{"class":215},"\"Worst offenders:",[205,1065,1066],{"class":1034},"\\n{",[205,1068,1069],{"class":244},"counts.head().to_string()",[205,1071,1041],{"class":1034},[205,1073,1074],{"class":215},"\"\n",[205,1076,1077],{"class":207,"line":388},[205,1078,1079],{"class":244},"        )\n",[205,1081,1082,1084],{"class":207,"line":413},[205,1083,631],{"class":240},[205,1085,1086],{"class":333}," True\n",[205,1088,1089],{"class":207,"line":437},[205,1090,258],{"emptyLinePlaceholder":257},[205,1092,1093,1096,1098,1100,1102],{"class":207,"line":461},[205,1094,1095],{"class":244},"check_unique(system, ",[205,1097,678],{"class":215},[205,1099,279],{"class":244},[205,1101,360],{"class":215},[205,1103,375],{"class":244},[205,1105,1106,1109,1111,1113,1115],{"class":207,"line":672},[205,1107,1108],{"class":244},"check_unique(ledger, ",[205,1110,678],{"class":215},[205,1112,279],{"class":244},[205,1114,466],{"class":215},[205,1116,375],{"class":244},[10,1118,1119],{},"Raising is usually right — a duplicated key in a file that should have unique ones is itself a finding. Where duplicates are legitimate, aggregate before comparing, or extend the key with the column that distinguishes them.",[190,1121,1123],{"id":1122},"step-3-the-reconciliation","Step 3 — The reconciliation",[10,1125,1126],{},"One outer merge answers both directions:",[195,1128,1130],{"className":231,"code":1129,"language":233,"meta":200,"style":200},"import pandas as pd\n\nmerged = system.merge(\n    ledger, on=\"key\", how=\"outer\", suffixes=(\"_system\", \"_ledger\"),\n    indicator=True,\n)\n\nonly_system = merged.loc[merged[\"_merge\"] == \"left_only\"]\nonly_ledger = merged.loc[merged[\"_merge\"] == \"right_only\"]\nin_both = merged.loc[merged[\"_merge\"] == \"both\"]\n\nprint(f\"only in system: {len(only_system)}\")\nprint(f\"only in ledger: {len(only_ledger)}\")\nprint(f\"in both:        {len(in_both)}\")\n",[202,1131,1132,1142,1146,1156,1198,1210,1214,1218,1242,1262,1282,1286,1313,1337],{"__ignoreMap":200},[205,1133,1134,1136,1138,1140],{"class":207,"line":208},[205,1135,241],{"class":240},[205,1137,245],{"class":244},[205,1139,248],{"class":240},[205,1141,251],{"class":244},[205,1143,1144],{"class":207,"line":254},[205,1145,258],{"emptyLinePlaceholder":257},[205,1147,1148,1151,1153],{"class":207,"line":261},[205,1149,1150],{"class":244},"merged ",[205,1152,369],{"class":240},[205,1154,1155],{"class":244}," system.merge(\n",[205,1157,1158,1161,1164,1166,1168,1170,1173,1175,1178,1180,1183,1185,1187,1190,1192,1195],{"class":207,"line":267},[205,1159,1160],{"class":244},"    ledger, ",[205,1162,1163],{"class":365},"on",[205,1165,369],{"class":240},[205,1167,678],{"class":215},[205,1169,279],{"class":244},[205,1171,1172],{"class":365},"how",[205,1174,369],{"class":240},[205,1176,1177],{"class":215},"\"outer\"",[205,1179,279],{"class":244},[205,1181,1182],{"class":365},"suffixes",[205,1184,369],{"class":240},[205,1186,800],{"class":244},[205,1188,1189],{"class":215},"\"_system\"",[205,1191,279],{"class":244},[205,1193,1194],{"class":215},"\"_ledger\"",[205,1196,1197],{"class":244},"),\n",[205,1199,1200,1203,1205,1207],{"class":207,"line":298},[205,1201,1202],{"class":365},"    indicator",[205,1204,369],{"class":240},[205,1206,623],{"class":333},[205,1208,1209],{"class":244},",\n",[205,1211,1212],{"class":207,"line":325},[205,1213,375],{"class":244},[205,1215,1216],{"class":207,"line":354},[205,1217,258],{"emptyLinePlaceholder":257},[205,1219,1220,1223,1225,1228,1231,1233,1236,1239],{"class":207,"line":378},[205,1221,1222],{"class":244},"only_system ",[205,1224,369],{"class":240},[205,1226,1227],{"class":244}," merged.loc[merged[",[205,1229,1230],{"class":215},"\"_merge\"",[205,1232,681],{"class":244},[205,1234,1235],{"class":240},"==",[205,1237,1238],{"class":215}," \"left_only\"",[205,1240,1241],{"class":244},"]\n",[205,1243,1244,1247,1249,1251,1253,1255,1257,1260],{"class":207,"line":383},[205,1245,1246],{"class":244},"only_ledger ",[205,1248,369],{"class":240},[205,1250,1227],{"class":244},[205,1252,1230],{"class":215},[205,1254,681],{"class":244},[205,1256,1235],{"class":240},[205,1258,1259],{"class":215}," \"right_only\"",[205,1261,1241],{"class":244},[205,1263,1264,1267,1269,1271,1273,1275,1277,1280],{"class":207,"line":388},[205,1265,1266],{"class":244},"in_both ",[205,1268,369],{"class":240},[205,1270,1227],{"class":244},[205,1272,1230],{"class":215},[205,1274,681],{"class":244},[205,1276,1235],{"class":240},[205,1278,1279],{"class":215}," \"both\"",[205,1281,1241],{"class":244},[205,1283,1284],{"class":207,"line":413},[205,1285,258],{"emptyLinePlaceholder":257},[205,1287,1288,1291,1293,1296,1299,1301,1304,1307,1309,1311],{"class":207,"line":437},[205,1289,1290],{"class":333},"print",[205,1292,800],{"class":244},[205,1294,1295],{"class":240},"f",[205,1297,1298],{"class":215},"\"only in system: ",[205,1300,1035],{"class":1034},[205,1302,1303],{"class":333},"len",[205,1305,1306],{"class":244},"(only_system)",[205,1308,1041],{"class":1034},[205,1310,602],{"class":215},[205,1312,375],{"class":244},[205,1314,1315,1317,1319,1321,1324,1326,1328,1331,1333,1335],{"class":207,"line":461},[205,1316,1290],{"class":333},[205,1318,800],{"class":244},[205,1320,1295],{"class":240},[205,1322,1323],{"class":215},"\"only in ledger: ",[205,1325,1035],{"class":1034},[205,1327,1303],{"class":333},[205,1329,1330],{"class":244},"(only_ledger)",[205,1332,1041],{"class":1034},[205,1334,602],{"class":215},[205,1336,375],{"class":244},[205,1338,1339,1341,1343,1345,1348,1350,1352,1355,1357,1359],{"class":207,"line":672},[205,1340,1290],{"class":333},[205,1342,800],{"class":244},[205,1344,1295],{"class":240},[205,1346,1347],{"class":215},"\"in both:        ",[205,1349,1035],{"class":1034},[205,1351,1303],{"class":333},[205,1353,1354],{"class":244},"(in_both)",[205,1356,1041],{"class":1034},[205,1358,602],{"class":215},[205,1360,375],{"class":244},[10,1362,1363,1364,1367,1368,1372],{},"Two mistakes to avoid. Using ",[202,1365,1366],{},"how=\"left\""," answers only one direction — the rows added on the other side stay invisible, which is exactly the case where a total is too high rather than too low. And comparing raw sets of order IDs rather than merging loses every other column, so you know ",[1369,1370,1371],"em",{},"that"," a row is missing but not what it contained.",[190,1374,1376],{"id":1375},"step-4-find-the-rows-that-differ","Step 4 — Find the rows that differ",[10,1378,1379,1380,1383],{},"The ",[202,1381,1382],{},"both"," group is where the subtler problems live: matching keys, different values.",[20,1385,30,1391,30,1394,30,1397,30,1400,30,1403,30,1408,30,1412,30,1417,30,1421,30,1425,30,1428,30,1433,30,1436,30,1439,30,1442,30,1445,30,1448,30,1453,30,1456,30,1459,30,1463],{"viewBox":1386,"role":23,"ariaLabel":1387,"ariaLabelledBy":1388,"xmlns":28,"style":29},"0 0 800 232","Three reconciliation outcomes: added rows, removed rows, and matched rows whose values differ, with the changed group requiring a column-by-column comparison.",[1389,1390],"three-t","three-d",[32,1392,1393],{"id":1389},"Added, removed, and changed — three findings, not two",[36,1395,1396],{"id":1390},"The indicator merge gives added and removed directly. The third and often most important group is rows whose key matched but whose values differ, which requires comparing each value column pairwise after the merge. A reconciliation that reports only added and removed will show two files as agreeing when every shared row has a different amount.",[40,1398],{"x":42,"y":42,"width":43,"height":1399,"fill":45},"232",[40,1401],{"x":48,"y":863,"width":108,"height":1402,"rx":48,"fill":53,"stroke":54,"style":55},"172",[57,1404,1407],{"x":170,"y":1405,"style":1406},"58","font-size:12.5px;font-weight:700;fill:var(--brand-strong,#4338ca);text-anchor:middle","added",[57,1409,1411],{"x":170,"y":1410,"style":936},"88","in B, not in A",[57,1413,1416],{"x":170,"y":1414,"style":1415},"114","font-size:11px;fill:var(--muted,#5b6780);text-anchor:middle","_merge == \"right_only\"",[57,1418,1420],{"x":170,"y":1419,"style":936},"148","new records, or records",[57,1422,1424],{"x":170,"y":1423,"style":936},"170","A has not yet received",[40,1426],{"x":1427,"y":863,"width":108,"height":1402,"rx":48,"fill":153,"stroke":154,"style":55},"276",[57,1429,1432],{"x":1430,"y":1405,"style":1431},"400","font-size:12.5px;font-weight:700;fill:var(--accent-ink,#be185d);text-anchor:middle","removed",[57,1434,1435],{"x":1430,"y":1410,"style":936},"in A, not in B",[57,1437,1438],{"x":1430,"y":1414,"style":1415},"_merge == \"left_only\"",[57,1440,1441],{"x":1430,"y":1419,"style":936},"deleted, or dropped by",[57,1443,1444],{"x":1430,"y":1423,"style":936},"a filter somewhere",[40,1446],{"x":1447,"y":863,"width":108,"height":1402,"rx":48,"fill":110,"stroke":111,"style":55},"538",[57,1449,1452],{"x":1450,"y":1405,"style":1451},"662","font-size:12.5px;font-weight:700;fill:var(--gold-ink,#7a4e06);text-anchor:middle","changed",[57,1454,1455],{"x":1450,"y":1410,"style":936},"key matches, values do not",[57,1457,1458],{"x":1450,"y":1414,"style":1415},"needs a pairwise compare",[57,1460,1462],{"x":1450,"y":1419,"style":1461},"font-size:11px;font-weight:700;fill:var(--gold-ink,#7a4e06);text-anchor:middle","the group most often missed",[57,1464,1465],{"x":1450,"y":1423,"style":1415},"and the one that moves totals",[195,1467,1469],{"className":231,"code":1468,"language":233,"meta":200,"style":200},"import numpy as np\nimport pandas as pd\n\ndef find_changes(merged, columns, tolerance=0.005):\n    \"\"\"Rows present in both files whose compared values differ.\"\"\"\n    both = merged.loc[merged[\"_merge\"] == \"both\"].copy()\n    differs = pd.Series(False, index=both.index)\n    details = {}\n\n    for name in columns:\n        left, right = f\"{name}_system\", f\"{name}_ledger\"\n        if left not in both or right not in both:\n            continue\n\n        if pd.api.types.is_numeric_dtype(both[left]):\n            gap = (both[left] - both[right]).abs()\n            changed = gap > tolerance\n            details[f\"{name}_delta\"] = np.where(changed, both[left] - both[right],\n                                                np.nan)\n        else:\n            changed = both[left].astype(\"string\").fillna(\"\") != \\\n                      both[right].astype(\"string\").fillna(\"\")\n\n        differs |= changed\n        details[f\"{name}_changed\"] = changed\n\n    result = both.loc[differs].assign(\n        **{k: pd.Series(v, index=both.index)[differs] for k, v in details.items()}\n    )\n    return result\n\nchanges = find_changes(merged, [\"region\", \"revenue\"])\nprint(f\"{len(changes)} row(s) differ\")\n",[202,1470,1471,1483,1493,1497,1515,1520,1540,1561,1571,1575,1589,1626,1656,1661,1665,1672,1689,1706,1737,1743,1752,1779,1793,1798,1810,1835,1840,1851,1878,1884,1892,1897,1918],{"__ignoreMap":200},[205,1472,1473,1475,1478,1480],{"class":207,"line":208},[205,1474,241],{"class":240},[205,1476,1477],{"class":244}," numpy ",[205,1479,248],{"class":240},[205,1481,1482],{"class":244}," np\n",[205,1484,1485,1487,1489,1491],{"class":207,"line":254},[205,1486,241],{"class":240},[205,1488,245],{"class":244},[205,1490,248],{"class":240},[205,1492,251],{"class":244},[205,1494,1495],{"class":207,"line":261},[205,1496,258],{"emptyLinePlaceholder":257},[205,1498,1499,1501,1504,1507,1509,1512],{"class":207,"line":267},[205,1500,528],{"class":240},[205,1502,1503],{"class":531}," find_changes",[205,1505,1506],{"class":244},"(merged, columns, tolerance",[205,1508,369],{"class":240},[205,1510,1511],{"class":333},"0.005",[205,1513,1514],{"class":244},"):\n",[205,1516,1517],{"class":207,"line":298},[205,1518,1519],{"class":215},"    \"\"\"Rows present in both files whose compared values differ.\"\"\"\n",[205,1521,1522,1525,1527,1529,1531,1533,1535,1537],{"class":207,"line":325},[205,1523,1524],{"class":244},"    both ",[205,1526,369],{"class":240},[205,1528,1227],{"class":244},[205,1530,1230],{"class":215},[205,1532,681],{"class":244},[205,1534,1235],{"class":240},[205,1536,1279],{"class":215},[205,1538,1539],{"class":244},"].copy()\n",[205,1541,1542,1545,1547,1550,1552,1554,1556,1558],{"class":207,"line":354},[205,1543,1544],{"class":244},"    differs ",[205,1546,369],{"class":240},[205,1548,1549],{"class":244}," pd.Series(",[205,1551,372],{"class":333},[205,1553,279],{"class":244},[205,1555,366],{"class":365},[205,1557,369],{"class":240},[205,1559,1560],{"class":244},"both.index)\n",[205,1562,1563,1566,1568],{"class":207,"line":378},[205,1564,1565],{"class":244},"    details ",[205,1567,369],{"class":240},[205,1569,1570],{"class":244}," {}\n",[205,1572,1573],{"class":207,"line":383},[205,1574,258],{"emptyLinePlaceholder":257},[205,1576,1577,1580,1583,1586],{"class":207,"line":388},[205,1578,1579],{"class":240},"    for",[205,1581,1582],{"class":244}," name ",[205,1584,1585],{"class":240},"in",[205,1587,1588],{"class":244}," columns:\n",[205,1590,1591,1594,1596,1599,1601,1603,1606,1608,1611,1613,1615,1617,1619,1621,1623],{"class":207,"line":413},[205,1592,1593],{"class":244},"        left, right ",[205,1595,369],{"class":240},[205,1597,1598],{"class":240}," f",[205,1600,602],{"class":215},[205,1602,1035],{"class":1034},[205,1604,1605],{"class":244},"name",[205,1607,1041],{"class":1034},[205,1609,1610],{"class":215},"_system\"",[205,1612,279],{"class":244},[205,1614,1295],{"class":240},[205,1616,602],{"class":215},[205,1618,1035],{"class":1034},[205,1620,1605],{"class":244},[205,1622,1041],{"class":1034},[205,1624,1625],{"class":215},"_ledger\"\n",[205,1627,1628,1631,1634,1637,1640,1643,1646,1649,1651,1653],{"class":207,"line":437},[205,1629,1630],{"class":240},"        if",[205,1632,1633],{"class":244}," left ",[205,1635,1636],{"class":240},"not",[205,1638,1639],{"class":240}," in",[205,1641,1642],{"class":244}," both ",[205,1644,1645],{"class":240},"or",[205,1647,1648],{"class":244}," right ",[205,1650,1636],{"class":240},[205,1652,1639],{"class":240},[205,1654,1655],{"class":244}," both:\n",[205,1657,1658],{"class":207,"line":461},[205,1659,1660],{"class":240},"            continue\n",[205,1662,1663],{"class":207,"line":672},[205,1664,258],{"emptyLinePlaceholder":257},[205,1666,1667,1669],{"class":207,"line":695},[205,1668,1630],{"class":240},[205,1670,1671],{"class":244}," pd.api.types.is_numeric_dtype(both[left]):\n",[205,1673,1675,1678,1680,1683,1686],{"class":207,"line":1674},16,[205,1676,1677],{"class":244},"            gap ",[205,1679,369],{"class":240},[205,1681,1682],{"class":244}," (both[left] ",[205,1684,1685],{"class":240},"-",[205,1687,1688],{"class":244}," both[right]).abs()\n",[205,1690,1692,1695,1697,1700,1703],{"class":207,"line":1691},17,[205,1693,1694],{"class":244},"            changed ",[205,1696,369],{"class":240},[205,1698,1699],{"class":244}," gap ",[205,1701,1702],{"class":240},">",[205,1704,1705],{"class":244}," tolerance\n",[205,1707,1709,1712,1714,1716,1718,1720,1722,1725,1727,1729,1732,1734],{"class":207,"line":1708},18,[205,1710,1711],{"class":244},"            details[",[205,1713,1295],{"class":240},[205,1715,602],{"class":215},[205,1717,1035],{"class":1034},[205,1719,1605],{"class":244},[205,1721,1041],{"class":1034},[205,1723,1724],{"class":215},"_delta\"",[205,1726,681],{"class":244},[205,1728,369],{"class":240},[205,1730,1731],{"class":244}," np.where(changed, both[left] ",[205,1733,1685],{"class":240},[205,1735,1736],{"class":244}," both[right],\n",[205,1738,1740],{"class":207,"line":1739},19,[205,1741,1742],{"class":244},"                                                np.nan)\n",[205,1744,1746,1749],{"class":207,"line":1745},20,[205,1747,1748],{"class":240},"        else",[205,1750,1751],{"class":244},":\n",[205,1753,1755,1757,1759,1762,1764,1767,1770,1773,1776],{"class":207,"line":1754},21,[205,1756,1694],{"class":244},[205,1758,369],{"class":240},[205,1760,1761],{"class":244}," both[left].astype(",[205,1763,553],{"class":215},[205,1765,1766],{"class":244},").fillna(",[205,1768,1769],{"class":215},"\"\"",[205,1771,1772],{"class":244},") ",[205,1774,1775],{"class":240},"!=",[205,1777,1778],{"class":244}," \\\n",[205,1780,1782,1785,1787,1789,1791],{"class":207,"line":1781},22,[205,1783,1784],{"class":244},"                      both[right].astype(",[205,1786,553],{"class":215},[205,1788,1766],{"class":244},[205,1790,1769],{"class":215},[205,1792,375],{"class":244},[205,1794,1796],{"class":207,"line":1795},23,[205,1797,258],{"emptyLinePlaceholder":257},[205,1799,1801,1804,1807],{"class":207,"line":1800},24,[205,1802,1803],{"class":244},"        differs ",[205,1805,1806],{"class":240},"|=",[205,1808,1809],{"class":244}," changed\n",[205,1811,1813,1816,1818,1820,1822,1824,1826,1829,1831,1833],{"class":207,"line":1812},25,[205,1814,1815],{"class":244},"        details[",[205,1817,1295],{"class":240},[205,1819,602],{"class":215},[205,1821,1035],{"class":1034},[205,1823,1605],{"class":244},[205,1825,1041],{"class":1034},[205,1827,1828],{"class":215},"_changed\"",[205,1830,681],{"class":244},[205,1832,369],{"class":240},[205,1834,1809],{"class":244},[205,1836,1838],{"class":207,"line":1837},26,[205,1839,258],{"emptyLinePlaceholder":257},[205,1841,1843,1846,1848],{"class":207,"line":1842},27,[205,1844,1845],{"class":244},"    result ",[205,1847,369],{"class":240},[205,1849,1850],{"class":244}," both.loc[differs].assign(\n",[205,1852,1854,1857,1860,1862,1864,1867,1870,1873,1875],{"class":207,"line":1853},28,[205,1855,1856],{"class":240},"        **",[205,1858,1859],{"class":244},"{k: pd.Series(v, ",[205,1861,366],{"class":365},[205,1863,369],{"class":240},[205,1865,1866],{"class":244},"both.index)[differs] ",[205,1868,1869],{"class":240},"for",[205,1871,1872],{"class":244}," k, v ",[205,1874,1585],{"class":240},[205,1876,1877],{"class":244}," details.items()}\n",[205,1879,1881],{"class":207,"line":1880},29,[205,1882,1883],{"class":244},"    )\n",[205,1885,1887,1889],{"class":207,"line":1886},30,[205,1888,631],{"class":240},[205,1890,1891],{"class":244}," result\n",[205,1893,1895],{"class":207,"line":1894},31,[205,1896,258],{"emptyLinePlaceholder":257},[205,1898,1900,1903,1905,1908,1911,1913,1916],{"class":207,"line":1899},32,[205,1901,1902],{"class":244},"changes ",[205,1904,369],{"class":240},[205,1906,1907],{"class":244}," find_changes(merged, [",[205,1909,1910],{"class":215},"\"region\"",[205,1912,279],{"class":244},[205,1914,1915],{"class":215},"\"revenue\"",[205,1917,692],{"class":244},[205,1919,1921,1923,1925,1927,1929,1931,1933,1936,1938,1941],{"class":207,"line":1920},33,[205,1922,1290],{"class":333},[205,1924,800],{"class":244},[205,1926,1295],{"class":240},[205,1928,602],{"class":215},[205,1930,1035],{"class":1034},[205,1932,1303],{"class":333},[205,1934,1935],{"class":244},"(changes)",[205,1937,1041],{"class":1034},[205,1939,1940],{"class":215}," row(s) differ\"",[205,1942,375],{"class":244},[10,1944,1945],{},"The numeric tolerance is not optional. Floating-point values that round-tripped through Excel differ in the fifteenth decimal place, and an exact comparison reports every single row as changed — which is worse than reporting none, because it buries the real differences.",[190,1947,1949],{"id":1948},"step-5-write-an-exceptions-workbook","Step 5 — Write an exceptions workbook",[10,1951,1952],{},"Three sheets and a summary give a colleague everything needed to act:",[195,1954,1956],{"className":231,"code":1955,"language":233,"meta":200,"style":200},"import pandas as pd\n\ndef reconciliation_report(only_left, only_right, changes, dest,\n                          left_name=\"system\", right_name=\"ledger\"):\n    \"\"\"Write a three-way reconciliation as a formatted workbook.\"\"\"\n    summary = pd.DataFrame({\n        \"finding\": [f\"only in {left_name}\", f\"only in {right_name}\", \"changed\"],\n        \"rows\": [len(only_left), len(only_right), len(changes)],\n    })\n\n    with pd.ExcelWriter(dest, engine=\"xlsxwriter\") as writer:\n        summary.to_excel(writer, sheet_name=\"Summary\", index=False)\n        only_left.to_excel(writer, sheet_name=f\"Only in {left_name}\"[:31],\n                           index=False)\n        only_right.to_excel(writer, sheet_name=f\"Only in {right_name}\"[:31],\n                            index=False)\n        changes.to_excel(writer, sheet_name=\"Changed\", index=False)\n\n        book = writer.book\n        header = book.add_format({\"bold\": True, \"bg_color\": \"#EEF2FF\",\n                                  \"border\": 1})\n        for name, frame in [(\"Summary\", summary),\n                            (f\"Only in {left_name}\"[:31], only_left),\n                            (f\"Only in {right_name}\"[:31], only_right),\n                            (\"Changed\", changes)]:\n            sheet = writer.sheets[name]\n            for position, column in enumerate(frame.columns):\n                sheet.write(0, position, str(column), header)\n                sheet.set_column(position, position, 18)\n            sheet.freeze_panes(1, 0)\n            if len(frame):\n                sheet.autofilter(0, 0, len(frame), len(frame.columns) - 1)\n\n    return dest\n\nreconciliation_report(only_system, only_ledger, changes, \"reconciliation.xlsx\")\n",[202,1957,1958,1968,1972,1982,2002,2007,2017,2060,2082,2087,2091,2114,2137,2167,2178,2205,2216,2238,2242,2252,2281,2294,2312,2336,2359,2368,2378,2394,2409,2419,2432,2443,2473,2477,2485,2490],{"__ignoreMap":200},[205,1959,1960,1962,1964,1966],{"class":207,"line":208},[205,1961,241],{"class":240},[205,1963,245],{"class":244},[205,1965,248],{"class":240},[205,1967,251],{"class":244},[205,1969,1970],{"class":207,"line":254},[205,1971,258],{"emptyLinePlaceholder":257},[205,1973,1974,1976,1979],{"class":207,"line":261},[205,1975,528],{"class":240},[205,1977,1978],{"class":531}," reconciliation_report",[205,1980,1981],{"class":244},"(only_left, only_right, changes, dest,\n",[205,1983,1984,1987,1989,1992,1995,1997,2000],{"class":207,"line":267},[205,1985,1986],{"class":244},"                          left_name",[205,1988,369],{"class":240},[205,1990,1991],{"class":215},"\"system\"",[205,1993,1994],{"class":244},", right_name",[205,1996,369],{"class":240},[205,1998,1999],{"class":215},"\"ledger\"",[205,2001,1514],{"class":244},[205,2003,2004],{"class":207,"line":298},[205,2005,2006],{"class":215},"    \"\"\"Write a three-way reconciliation as a formatted workbook.\"\"\"\n",[205,2008,2009,2012,2014],{"class":207,"line":325},[205,2010,2011],{"class":244},"    summary ",[205,2013,369],{"class":240},[205,2015,2016],{"class":244}," pd.DataFrame({\n",[205,2018,2019,2022,2024,2026,2029,2031,2034,2036,2038,2040,2042,2044,2046,2049,2051,2053,2055,2058],{"class":207,"line":354},[205,2020,2021],{"class":215},"        \"finding\"",[205,2023,273],{"class":244},[205,2025,1295],{"class":240},[205,2027,2028],{"class":215},"\"only in ",[205,2030,1035],{"class":1034},[205,2032,2033],{"class":244},"left_name",[205,2035,1041],{"class":1034},[205,2037,602],{"class":215},[205,2039,279],{"class":244},[205,2041,1295],{"class":240},[205,2043,2028],{"class":215},[205,2045,1035],{"class":1034},[205,2047,2048],{"class":244},"right_name",[205,2050,1041],{"class":1034},[205,2052,602],{"class":215},[205,2054,279],{"class":244},[205,2056,2057],{"class":215},"\"changed\"",[205,2059,295],{"class":244},[205,2061,2062,2065,2067,2069,2072,2074,2077,2079],{"class":207,"line":378},[205,2063,2064],{"class":215},"        \"rows\"",[205,2066,273],{"class":244},[205,2068,1303],{"class":333},[205,2070,2071],{"class":244},"(only_left), ",[205,2073,1303],{"class":333},[205,2075,2076],{"class":244},"(only_right), ",[205,2078,1303],{"class":333},[205,2080,2081],{"class":244},"(changes)],\n",[205,2083,2084],{"class":207,"line":383},[205,2085,2086],{"class":244},"    })\n",[205,2088,2089],{"class":207,"line":388},[205,2090,258],{"emptyLinePlaceholder":257},[205,2092,2093,2096,2099,2102,2104,2107,2109,2111],{"class":207,"line":413},[205,2094,2095],{"class":240},"    with",[205,2097,2098],{"class":244}," pd.ExcelWriter(dest, ",[205,2100,2101],{"class":365},"engine",[205,2103,369],{"class":240},[205,2105,2106],{"class":215},"\"xlsxwriter\"",[205,2108,1772],{"class":244},[205,2110,248],{"class":240},[205,2112,2113],{"class":244}," writer:\n",[205,2115,2116,2119,2122,2124,2127,2129,2131,2133,2135],{"class":207,"line":437},[205,2117,2118],{"class":244},"        summary.to_excel(writer, ",[205,2120,2121],{"class":365},"sheet_name",[205,2123,369],{"class":240},[205,2125,2126],{"class":215},"\"Summary\"",[205,2128,279],{"class":244},[205,2130,366],{"class":365},[205,2132,369],{"class":240},[205,2134,372],{"class":333},[205,2136,375],{"class":244},[205,2138,2139,2142,2144,2146,2148,2151,2153,2155,2157,2159,2162,2165],{"class":207,"line":461},[205,2140,2141],{"class":244},"        only_left.to_excel(writer, ",[205,2143,2121],{"class":365},[205,2145,369],{"class":240},[205,2147,1295],{"class":240},[205,2149,2150],{"class":215},"\"Only in ",[205,2152,1035],{"class":1034},[205,2154,2033],{"class":244},[205,2156,1041],{"class":1034},[205,2158,602],{"class":215},[205,2160,2161],{"class":244},"[:",[205,2163,2164],{"class":333},"31",[205,2166,295],{"class":244},[205,2168,2169,2172,2174,2176],{"class":207,"line":672},[205,2170,2171],{"class":365},"                           index",[205,2173,369],{"class":240},[205,2175,372],{"class":333},[205,2177,375],{"class":244},[205,2179,2180,2183,2185,2187,2189,2191,2193,2195,2197,2199,2201,2203],{"class":207,"line":695},[205,2181,2182],{"class":244},"        only_right.to_excel(writer, ",[205,2184,2121],{"class":365},[205,2186,369],{"class":240},[205,2188,1295],{"class":240},[205,2190,2150],{"class":215},[205,2192,1035],{"class":1034},[205,2194,2048],{"class":244},[205,2196,1041],{"class":1034},[205,2198,602],{"class":215},[205,2200,2161],{"class":244},[205,2202,2164],{"class":333},[205,2204,295],{"class":244},[205,2206,2207,2210,2212,2214],{"class":207,"line":1674},[205,2208,2209],{"class":365},"                            index",[205,2211,369],{"class":240},[205,2213,372],{"class":333},[205,2215,375],{"class":244},[205,2217,2218,2221,2223,2225,2228,2230,2232,2234,2236],{"class":207,"line":1691},[205,2219,2220],{"class":244},"        changes.to_excel(writer, ",[205,2222,2121],{"class":365},[205,2224,369],{"class":240},[205,2226,2227],{"class":215},"\"Changed\"",[205,2229,279],{"class":244},[205,2231,366],{"class":365},[205,2233,369],{"class":240},[205,2235,372],{"class":333},[205,2237,375],{"class":244},[205,2239,2240],{"class":207,"line":1708},[205,2241,258],{"emptyLinePlaceholder":257},[205,2243,2244,2247,2249],{"class":207,"line":1739},[205,2245,2246],{"class":244},"        book ",[205,2248,369],{"class":240},[205,2250,2251],{"class":244}," writer.book\n",[205,2253,2254,2257,2259,2262,2265,2267,2269,2271,2274,2276,2279],{"class":207,"line":1745},[205,2255,2256],{"class":244},"        header ",[205,2258,369],{"class":240},[205,2260,2261],{"class":244}," book.add_format({",[205,2263,2264],{"class":215},"\"bold\"",[205,2266,1044],{"class":244},[205,2268,623],{"class":333},[205,2270,279],{"class":244},[205,2272,2273],{"class":215},"\"bg_color\"",[205,2275,1044],{"class":244},[205,2277,2278],{"class":215},"\"#EEF2FF\"",[205,2280,1209],{"class":244},[205,2282,2283,2286,2288,2291],{"class":207,"line":1754},[205,2284,2285],{"class":215},"                                  \"border\"",[205,2287,1044],{"class":244},[205,2289,2290],{"class":333},"1",[205,2292,2293],{"class":244},"})\n",[205,2295,2296,2299,2302,2304,2307,2309],{"class":207,"line":1781},[205,2297,2298],{"class":240},"        for",[205,2300,2301],{"class":244}," name, frame ",[205,2303,1585],{"class":240},[205,2305,2306],{"class":244}," [(",[205,2308,2126],{"class":215},[205,2310,2311],{"class":244},", summary),\n",[205,2313,2314,2317,2319,2321,2323,2325,2327,2329,2331,2333],{"class":207,"line":1795},[205,2315,2316],{"class":244},"                            (",[205,2318,1295],{"class":240},[205,2320,2150],{"class":215},[205,2322,1035],{"class":1034},[205,2324,2033],{"class":244},[205,2326,1041],{"class":1034},[205,2328,602],{"class":215},[205,2330,2161],{"class":244},[205,2332,2164],{"class":333},[205,2334,2335],{"class":244},"], only_left),\n",[205,2337,2338,2340,2342,2344,2346,2348,2350,2352,2354,2356],{"class":207,"line":1800},[205,2339,2316],{"class":244},[205,2341,1295],{"class":240},[205,2343,2150],{"class":215},[205,2345,1035],{"class":1034},[205,2347,2048],{"class":244},[205,2349,1041],{"class":1034},[205,2351,602],{"class":215},[205,2353,2161],{"class":244},[205,2355,2164],{"class":333},[205,2357,2358],{"class":244},"], only_right),\n",[205,2360,2361,2363,2365],{"class":207,"line":1812},[205,2362,2316],{"class":244},[205,2364,2227],{"class":215},[205,2366,2367],{"class":244},", changes)]:\n",[205,2369,2370,2373,2375],{"class":207,"line":1837},[205,2371,2372],{"class":244},"            sheet ",[205,2374,369],{"class":240},[205,2376,2377],{"class":244}," writer.sheets[name]\n",[205,2379,2380,2383,2386,2388,2391],{"class":207,"line":1842},[205,2381,2382],{"class":240},"            for",[205,2384,2385],{"class":244}," position, column ",[205,2387,1585],{"class":240},[205,2389,2390],{"class":333}," enumerate",[205,2392,2393],{"class":244},"(frame.columns):\n",[205,2395,2396,2399,2401,2404,2406],{"class":207,"line":1853},[205,2397,2398],{"class":244},"                sheet.write(",[205,2400,42],{"class":333},[205,2402,2403],{"class":244},", position, ",[205,2405,797],{"class":333},[205,2407,2408],{"class":244},"(column), header)\n",[205,2410,2411,2414,2417],{"class":207,"line":1880},[205,2412,2413],{"class":244},"                sheet.set_column(position, position, ",[205,2415,2416],{"class":333},"18",[205,2418,375],{"class":244},[205,2420,2421,2424,2426,2428,2430],{"class":207,"line":1886},[205,2422,2423],{"class":244},"            sheet.freeze_panes(",[205,2425,2290],{"class":333},[205,2427,279],{"class":244},[205,2429,42],{"class":333},[205,2431,375],{"class":244},[205,2433,2434,2437,2440],{"class":207,"line":1894},[205,2435,2436],{"class":240},"            if",[205,2438,2439],{"class":333}," len",[205,2441,2442],{"class":244},"(frame):\n",[205,2444,2445,2448,2450,2452,2454,2456,2458,2461,2463,2466,2468,2471],{"class":207,"line":1899},[205,2446,2447],{"class":244},"                sheet.autofilter(",[205,2449,42],{"class":333},[205,2451,279],{"class":244},[205,2453,42],{"class":333},[205,2455,279],{"class":244},[205,2457,1303],{"class":333},[205,2459,2460],{"class":244},"(frame), ",[205,2462,1303],{"class":333},[205,2464,2465],{"class":244},"(frame.columns) ",[205,2467,1685],{"class":240},[205,2469,2470],{"class":333}," 1",[205,2472,375],{"class":244},[205,2474,2475],{"class":207,"line":1920},[205,2476,258],{"emptyLinePlaceholder":257},[205,2478,2480,2482],{"class":207,"line":2479},34,[205,2481,631],{"class":240},[205,2483,2484],{"class":244}," dest\n",[205,2486,2488],{"class":207,"line":2487},35,[205,2489,258],{"emptyLinePlaceholder":257},[205,2491,2493,2496,2499],{"class":207,"line":2492},36,[205,2494,2495],{"class":244},"reconciliation_report(only_system, only_ledger, changes, ",[205,2497,2498],{"class":215},"\"reconciliation.xlsx\"",[205,2500,375],{"class":244},[10,2502,2503,2504,18],{},"Leading with a summary sheet matters: a reader wants the three counts before the detail, and a report that opens on eighteen rows of exceptions with no context invites the question \"out of how many?\". The multi-sheet mechanics are covered in ",[14,2505,2507],{"href":2506},"\u002Fautomating-reporting-workflows\u002Fbuilding-multi-sheet-excel-dashboards\u002Fadd-summary-sheet-to-excel-report-python\u002F","adding a summary sheet to an Excel report",[190,2509,2511],{"id":2510},"common-pitfalls-and-fixes","Common pitfalls and fixes",[2513,2514,2515,2531],"table",{},[2516,2517,2518],"thead",{},[2519,2520,2521,2525,2528],"tr",{},[2522,2523,2524],"th",{},"Symptom",[2522,2526,2527],{},"Cause",[2522,2529,2530],{},"Fix",[2532,2533,2534,2546,2557,2575,2586,2601,2612],"tbody",{},[2519,2535,2536,2540,2543],{},[2537,2538,2539],"td",{},"Rows report missing from both sides",[2537,2541,2542],{},"Key differs invisibly",[2537,2544,2545],{},"Normalise both keys before merging.",[2519,2547,2548,2551,2554],{},[2537,2549,2550],{},"Row count explodes after the merge",[2537,2552,2553],{},"Duplicated key on one side",[2537,2555,2556],{},"Check uniqueness; aggregate or extend the key.",[2519,2558,2559,2562,2566],{},[2537,2560,2561],{},"Only one direction reported",[2537,2563,2564],{},[202,2565,1366],{},[2537,2567,2568,2569,2572,2573,18],{},"Use ",[202,2570,2571],{},"how=\"outer\""," with ",[202,2574,122],{},[2519,2576,2577,2580,2583],{},[2537,2578,2579],{},"Every shared row reports as changed",[2537,2581,2582],{},"Exact float comparison",[2537,2584,2585],{},"Compare numerics with a tolerance.",[2519,2587,2588,2595,2598],{},[2537,2589,2590,2592,2593],{},[202,2591,717],{}," does not match ",[202,2594,721],{},[2537,2596,2597],{},"Number stored as text on one side",[2537,2599,2600],{},"Canonicalise numeric keys.",[2519,2602,2603,2606,2609],{},[2537,2604,2605],{},"Missing rows found but not their contents",[2537,2607,2608],{},"Compared sets of IDs, not frames",[2537,2610,2611],{},"Merge the frames, not the key columns.",[2519,2613,2614,2617,2620],{},[2537,2615,2616],{},"Report is unusable",[2537,2618,2619],{},"Raw dump with no summary",[2537,2621,2622],{},"Lead with counts, then the detail sheets.",[190,2624,2626],{"id":2625},"performance-and-scale-notes","Performance and scale notes",[10,2628,2629,2632],{},[202,2630,2631],{},"merge"," builds a hash index over the keys, so it is roughly linear and handles a few million rows comfortably. The costs sit around it.",[10,2634,2635,2639],{},[2636,2637,2638],"strong",{},"Read only what you compare."," A reconciliation on four columns has no reason to load forty:",[195,2641,2643],{"className":231,"code":2642,"language":233,"meta":200,"style":200},"COMPARE = [\"order_id\", \"region\", \"revenue\"]\nsystem = pd.read_excel(\"system.xlsx\", usecols=COMPARE)\nledger = pd.read_excel(\"ledger.xlsx\", usecols=COMPARE)\n",[202,2644,2645,2668,2689],{"__ignoreMap":200},[205,2646,2647,2650,2653,2656,2658,2660,2662,2664,2666],{"class":207,"line":208},[205,2648,2649],{"class":333},"COMPARE",[205,2651,2652],{"class":240}," =",[205,2654,2655],{"class":244}," [",[205,2657,689],{"class":215},[205,2659,279],{"class":244},[205,2661,1910],{"class":215},[205,2663,279],{"class":244},[205,2665,1915],{"class":215},[205,2667,1241],{"class":244},[205,2669,2670,2672,2674,2676,2678,2680,2683,2685,2687],{"class":207,"line":254},[205,2671,643],{"class":244},[205,2673,369],{"class":240},[205,2675,648],{"class":244},[205,2677,360],{"class":215},[205,2679,279],{"class":244},[205,2681,2682],{"class":365},"usecols",[205,2684,369],{"class":240},[205,2686,2649],{"class":333},[205,2688,375],{"class":244},[205,2690,2691,2693,2695,2697,2699,2701,2703,2705,2707],{"class":207,"line":261},[205,2692,657],{"class":244},[205,2694,369],{"class":240},[205,2696,648],{"class":244},[205,2698,466],{"class":215},[205,2700,279],{"class":244},[205,2702,2682],{"class":365},[205,2704,369],{"class":240},[205,2706,2649],{"class":333},[205,2708,375],{"class":244},[10,2710,2711,2714],{},[2636,2712,2713],{},"Normalise the distinct key values, not every row."," Where the key is low-cardinality — a product or account code — cleaning the unique values and mapping is far cheaper than cleaning a million strings.",[10,2716,2717,2723],{},[2636,2718,2568,2719,2722],{},[202,2720,2721],{},"validate"," to fail fast."," pandas will check the join cardinality for you, which is cheaper and clearer than discovering the row explosion afterwards:",[195,2725,2727],{"className":231,"code":2726,"language":233,"meta":200,"style":200},"merged = system.merge(\n    ledger, on=\"key\", how=\"outer\", indicator=True,\n    suffixes=(\"_system\", \"_ledger\"),\n    validate=\"one_to_one\",       # raises immediately if either side duplicates\n)\n",[202,2728,2729,2737,2766,2783,2799],{"__ignoreMap":200},[205,2730,2731,2733,2735],{"class":207,"line":208},[205,2732,1150],{"class":244},[205,2734,369],{"class":240},[205,2736,1155],{"class":244},[205,2738,2739,2741,2743,2745,2747,2749,2751,2753,2755,2757,2760,2762,2764],{"class":207,"line":254},[205,2740,1160],{"class":244},[205,2742,1163],{"class":365},[205,2744,369],{"class":240},[205,2746,678],{"class":215},[205,2748,279],{"class":244},[205,2750,1172],{"class":365},[205,2752,369],{"class":240},[205,2754,1177],{"class":215},[205,2756,279],{"class":244},[205,2758,2759],{"class":365},"indicator",[205,2761,369],{"class":240},[205,2763,623],{"class":333},[205,2765,1209],{"class":244},[205,2767,2768,2771,2773,2775,2777,2779,2781],{"class":207,"line":261},[205,2769,2770],{"class":365},"    suffixes",[205,2772,369],{"class":240},[205,2774,800],{"class":244},[205,2776,1189],{"class":215},[205,2778,279],{"class":244},[205,2780,1194],{"class":215},[205,2782,1197],{"class":244},[205,2784,2785,2788,2790,2793,2796],{"class":207,"line":267},[205,2786,2787],{"class":365},"    validate",[205,2789,369],{"class":240},[205,2791,2792],{"class":215},"\"one_to_one\"",[205,2794,2795],{"class":244},",       ",[205,2797,2798],{"class":587},"# raises immediately if either side duplicates\n",[205,2800,2801],{"class":207,"line":298},[205,2802,375],{"class":244},[10,2804,2805],{},"For genuinely large files, compare hashes rather than values. Hashing each row to a single digest turns a wide comparison into a one-column one, which both reads and merges faster:",[195,2807,2809],{"className":231,"code":2808,"language":233,"meta":200,"style":200},"import hashlib\nimport pandas as pd\n\ndef row_digest(df, columns):\n    joined = df[columns].astype(\"string\").fillna(\"\").agg(\"|\".join, axis=1)\n    return joined.map(lambda s: hashlib.md5(s.encode()).hexdigest())\n\nsystem[\"digest\"] = row_digest(system, [\"region\", \"revenue\"])\nledger[\"digest\"] = row_digest(ledger, [\"region\", \"revenue\"])\nchanged_keys = system.merge(ledger, on=\"key\")\nchanged_keys = changed_keys.loc[\n    changed_keys[\"digest_x\"] != changed_keys[\"digest_y\"], \"key\"\n]\n",[202,2810,2811,2818,2828,2832,2842,2876,2889,2893,2915,2936,2954,2963,2987],{"__ignoreMap":200},[205,2812,2813,2815],{"class":207,"line":208},[205,2814,241],{"class":240},[205,2816,2817],{"class":244}," hashlib\n",[205,2819,2820,2822,2824,2826],{"class":207,"line":254},[205,2821,241],{"class":240},[205,2823,245],{"class":244},[205,2825,248],{"class":240},[205,2827,251],{"class":244},[205,2829,2830],{"class":207,"line":261},[205,2831,258],{"emptyLinePlaceholder":257},[205,2833,2834,2836,2839],{"class":207,"line":267},[205,2835,528],{"class":240},[205,2837,2838],{"class":531}," row_digest",[205,2840,2841],{"class":244},"(df, columns):\n",[205,2843,2844,2847,2849,2852,2854,2856,2858,2861,2864,2867,2870,2872,2874],{"class":207,"line":298},[205,2845,2846],{"class":244},"    joined ",[205,2848,369],{"class":240},[205,2850,2851],{"class":244}," df[columns].astype(",[205,2853,553],{"class":215},[205,2855,1766],{"class":244},[205,2857,1769],{"class":215},[205,2859,2860],{"class":244},").agg(",[205,2862,2863],{"class":215},"\"|\"",[205,2865,2866],{"class":244},".join, ",[205,2868,2869],{"class":365},"axis",[205,2871,369],{"class":240},[205,2873,2290],{"class":333},[205,2875,375],{"class":244},[205,2877,2878,2880,2883,2886],{"class":207,"line":325},[205,2879,631],{"class":240},[205,2881,2882],{"class":244}," joined.map(",[205,2884,2885],{"class":240},"lambda",[205,2887,2888],{"class":244}," s: hashlib.md5(s.encode()).hexdigest())\n",[205,2890,2891],{"class":207,"line":354},[205,2892,258],{"emptyLinePlaceholder":257},[205,2894,2895,2897,2900,2902,2904,2907,2909,2911,2913],{"class":207,"line":378},[205,2896,675],{"class":244},[205,2898,2899],{"class":215},"\"digest\"",[205,2901,681],{"class":244},[205,2903,369],{"class":240},[205,2905,2906],{"class":244}," row_digest(system, [",[205,2908,1910],{"class":215},[205,2910,279],{"class":244},[205,2912,1915],{"class":215},[205,2914,692],{"class":244},[205,2916,2917,2919,2921,2923,2925,2928,2930,2932,2934],{"class":207,"line":383},[205,2918,698],{"class":244},[205,2920,2899],{"class":215},[205,2922,681],{"class":244},[205,2924,369],{"class":240},[205,2926,2927],{"class":244}," row_digest(ledger, [",[205,2929,1910],{"class":215},[205,2931,279],{"class":244},[205,2933,1915],{"class":215},[205,2935,692],{"class":244},[205,2937,2938,2941,2943,2946,2948,2950,2952],{"class":207,"line":388},[205,2939,2940],{"class":244},"changed_keys ",[205,2942,369],{"class":240},[205,2944,2945],{"class":244}," system.merge(ledger, ",[205,2947,1163],{"class":365},[205,2949,369],{"class":240},[205,2951,678],{"class":215},[205,2953,375],{"class":244},[205,2955,2956,2958,2960],{"class":207,"line":413},[205,2957,2940],{"class":244},[205,2959,369],{"class":240},[205,2961,2962],{"class":244}," changed_keys.loc[\n",[205,2964,2965,2968,2971,2973,2975,2978,2981,2984],{"class":207,"line":437},[205,2966,2967],{"class":244},"    changed_keys[",[205,2969,2970],{"class":215},"\"digest_x\"",[205,2972,681],{"class":244},[205,2974,1775],{"class":240},[205,2976,2977],{"class":244}," changed_keys[",[205,2979,2980],{"class":215},"\"digest_y\"",[205,2982,2983],{"class":244},"], ",[205,2985,2986],{"class":215},"\"key\"\n",[205,2988,2989],{"class":207,"line":461},[205,2990,1241],{"class":244},[10,2992,2993,2994,2997,2998,3002],{},"That identifies ",[1369,2995,2996],{},"which"," rows changed cheaply; fetch the details only for those. Where either file is too large to hold at all, the chunked reading approach in ",[14,2999,3001],{"href":3000},"\u002Fadvanced-data-transformation-and-cleaning\u002Fworking-with-large-excel-files-in-python\u002Fread-large-excel-file-in-chunks-with-pandas\u002F","reading large Excel files in chunks"," lets you build a key-to-digest mapping from one side and stream the other against it.",[190,3004,3006],{"id":3005},"conclusion","Conclusion",[10,3008,3009,3010,3012],{},"Reconciling two spreadsheets is one outer merge with ",[202,3011,122],{}," — surrounded by the work that makes its answer true. Normalise both keys so whitespace and formatting cannot manufacture differences, confirm the keys are unique before merging so the counts mean something, and compare numeric values with a tolerance so floating-point noise does not flag every row. Report all three findings, not two: added, removed, and the matched rows whose values changed. Then lead the workbook with a summary, because the counts are what a reader needs before the detail.",[190,3014,3016],{"id":3015},"frequently-asked-questions","Frequently asked questions",[10,3018,3019,3022,3023,718,3025,3027,3028,3031,3032,3034],{},[2636,3020,3021],{},"What is the quickest way to find rows in A that are not in B?","\nMerge them on the key with ",[202,3024,1366],{},[202,3026,122],{},", then keep the rows whose merge indicator is ",[202,3029,3030],{},"left_only",". It is one pass and it reports both directions when you use ",[202,3033,2571],{}," instead.",[10,3036,3037,3040],{},[2636,3038,3039],{},"Why do rows show as missing when I can see them in both files?","\nThe keys differ invisibly — trailing whitespace, a non-breaking space, different case, or a number stored as text on one side. Normalise both keys into a comparison column before merging.",[10,3042,3043,3046,3047,3049],{},[2636,3044,3045],{},"How do I compare on more than one column?","\nPass a list to the ",[202,3048,1163],{}," argument, or build a single composite key by joining the normalised parts with a separator that cannot appear in the values, such as a vertical bar.",[10,3051,3052,3055],{},[2636,3053,3054],{},"What if the key is not unique?","\nA merge on a duplicated key multiplies rows. Check for duplicates first and decide deliberately — aggregate them, keep the latest, or treat the duplication itself as the finding.",[10,3057,3058,3061],{},[2636,3059,3060],{},"Can I compare the values as well as the keys?","\nYes. Merge on the key with suffixes, then compare the value columns pairwise to classify each matched row as identical or changed. That turns a two-way difference into a three-way one: added, removed and changed.",[190,3063,3065],{"id":3064},"related","Related",[3067,3068,3069,3076,3083,3090,3096],"ul",{},[3070,3071,3072,3073,3075],"li",{},"Up to the parent: ",[14,3074,17],{"href":16}," — the join mechanics this builds on.",[3070,3077,3078,3082],{},[14,3079,3081],{"href":3080},"\u002Fadvanced-data-transformation-and-cleaning\u002Fvalidating-excel-data-with-python\u002Fcompare-two-excel-files-for-differences-with-python\u002F","Compare Two Excel Files for Differences with Python"," — cell-level comparison of two whole sheets.",[3070,3084,3085,3089],{},[14,3086,3088],{"href":3087},"\u002Fadvanced-data-transformation-and-cleaning\u002Fmerging-and-joining-excel-dataframes\u002Fconcatenate-excel-sheets-with-different-columns\u002F","Concatenate Excel Sheets with Different Columns"," — combining the files this compares.",[3070,3091,3092,3095],{},[14,3093,3094],{"href":838},"Strip Whitespace and Normalise Text Columns with pandas"," — the key normalisation this depends on.",[3070,3097,3098,3102],{},[14,3099,3101],{"href":3100},"\u002Fadvanced-data-transformation-and-cleaning\u002Fvalidating-excel-data-with-python\u002Ffind-duplicate-rows-in-excel-with-python\u002F","Find Duplicate Rows in Excel with Python"," — the uniqueness check before merging.",[3104,3105,3106],"style",{},"html pre.shiki code .sMTad, html code.shiki .sMTad{--shiki-default:#6F42C1;--shiki-dark:#FFB757}html pre.shiki code .srMev, html code.shiki .srMev{--shiki-default:#032F62;--shiki-dark:#ADDCFF}html .default .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html.dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html pre.shiki code .s-kum, html code.shiki .s-kum{--shiki-default:#D73A49;--shiki-dark:#FF9492}html pre.shiki code .skGVy, html code.shiki .skGVy{--shiki-default:#24292E;--shiki-dark:#F0F3F6}html pre.shiki code .sP0c6, html code.shiki .sP0c6{--shiki-default:#005CC5;--shiki-dark:#91CBFF}html pre.shiki code .sa561, html code.shiki .sa561{--shiki-default:#E36209;--shiki-dark:#FFB757}html pre.shiki code .s_Opv, html code.shiki .s_Opv{--shiki-default:#6F42C1;--shiki-dark:#DBB7FF}html pre.shiki code .s-wDw, html code.shiki .s-wDw{--shiki-default:#6A737D;--shiki-dark:#BDC4CC}html pre.shiki code .sSjpA, html code.shiki .sSjpA{--shiki-default:#005CC5;--shiki-dark:#FF9492}",{"title":200,"searchDepth":254,"depth":254,"links":3108},[3109,3110,3111,3112,3113,3114,3115,3116,3117,3118,3119],{"id":192,"depth":254,"text":193},{"id":493,"depth":254,"text":494},{"id":842,"depth":254,"text":843},{"id":1122,"depth":254,"text":1123},{"id":1375,"depth":254,"text":1376},{"id":1948,"depth":254,"text":1949},{"id":2510,"depth":254,"text":2511},{"id":2625,"depth":254,"text":2626},{"id":3005,"depth":254,"text":3006},{"id":3015,"depth":254,"text":3016},{"id":3064,"depth":254,"text":3065},"2026-08-15","Reconcile two spreadsheets in Python — an indicator merge for both directions, composite keys, near-matches from whitespace, and a formatted exceptions workbook.","md",[3124,3126,3128,3130,3132],{"q":3021,"a":3125},"Merge them on the key with how=\"left\" and indicator=True, then keep the rows whose merge indicator is left_only. It is one pass and it reports both directions when you use how=\"outer\" instead.",{"q":3039,"a":3127},"The keys differ invisibly — trailing whitespace, a non-breaking space, different case, or a number stored as text on one side. Normalise both keys into a comparison column before merging.",{"q":3045,"a":3129},"Pass a list to the on argument, or build a single composite key by joining the normalised parts with a separator that cannot appear in the values, such as a vertical bar.",{"q":3054,"a":3131},"A merge on a duplicated key multiplies rows. Check for duplicates first and decide deliberately — aggregate them, keep the latest, or treat the duplication itself as the finding.",{"q":3060,"a":3133},{"Yes":3134},{" Merge on the key with suffixes, then compare the value columns pairwise to classify each matched row as identical or changed":3135},{" That turns a two-way difference into a three-way one":3136},"added, removed and changed.",{},"\u002Fadvanced-data-transformation-and-cleaning\u002Fmerging-and-joining-excel-dataframes\u002Ffind-rows-in-one-excel-file-missing-from-another",{"title":3140,"description":3141},"Find Missing Rows Between Two Excel Files (pandas)","Compare two Excel files row by row with pandas: merge with indicator, both-direction differences, composite and normalised keys, duplicates, and an exceptions report.","find-rows-in-one-excel-file-missing-from-another","advanced-data-transformation-and-cleaning\u002Fmerging-and-joining-excel-dataframes\u002Ffind-rows-in-one-excel-file-missing-from-another\u002Findex","how-to","lnlY_YkCD2t5qlhwa8N18fMqWvbIHX9AJAv7fSDrLjw",[3147,3150],{"title":3088,"path":3148,"stem":3149,"children":-1},"\u002Fadvanced-data-transformation-and-cleaning\u002Fmerging-and-joining-excel-dataframes\u002Fconcatenate-excel-sheets-with-different-columns","advanced-data-transformation-and-cleaning\u002Fmerging-and-joining-excel-dataframes\u002Fconcatenate-excel-sheets-with-different-columns\u002Findex",{"title":3151,"path":3152,"stem":3153,"children":-1},"Merge Two Excel Files on a Common Column in Python","\u002Fadvanced-data-transformation-and-cleaning\u002Fmerging-and-joining-excel-dataframes\u002Fmerge-two-excel-files-on-common-column-python","advanced-data-transformation-and-cleaning\u002Fmerging-and-joining-excel-dataframes\u002Fmerge-two-excel-files-on-common-column-python\u002Findex",1786800028629]