[{"data":1,"prerenderedAt":2550},["ShallowReactive",2],{"doc:\u002Fgetting-started-with-python-excel-automation\u002Freading-excel-files-with-pandas\u002Fskip-rows-and-set-header-when-reading-excel-with-pandas":3,"surround:\u002Fgetting-started-with-python-excel-automation\u002Freading-excel-files-with-pandas\u002Fskip-rows-and-set-header-when-reading-excel-with-pandas":2541},{"id":4,"title":5,"body":6,"dateModified":2516,"datePublished":2516,"description":2517,"extension":2518,"faq":2519,"meta":2532,"navigation":238,"path":2533,"seo":2534,"slug":2537,"stem":2538,"type":2539,"__hash__":2540},"docs\u002Fgetting-started-with-python-excel-automation\u002Freading-excel-files-with-pandas\u002Fskip-rows-and-set-header-when-reading-excel-with-pandas\u002Findex.md","Skip Rows and Set the Header When Reading Excel with pandas",{"type":7,"value":8,"toc":2503},"minimark",[9,28,174,179,207,210,474,478,481,578,588,592,595,647,671,751,768,847,851,967,983,1048,1058,1229,1242,1246,1252,1338,1711,1725,1729,1736,1769,1775,1863,1866,1869,1980,1984,2129,2133,2146,2149,2334,2341,2350,2366,2370,2385,2389,2409,2421,2430,2441,2454,2458,2499],[10,11,12,13,17,18,21,22,27],"p",{},"Excel files made by people rarely start with the data. There is a company banner, a title, an \"as at\" date, a blank row, then the actual column headings — and pandas, reading from row zero, dutifully names your columns ",[14,15,16],"code",{},"Unnamed: 0"," through ",[14,19,20],{},"Unnamed: 7",". The fix is two arguments, but knowing which one to reach for, and what to do when the preamble changes height every month, is what makes an import robust. This guide covers both. It builds on the basics in ",[23,24,26],"a",{"href":25},"\u002Fgetting-started-with-python-excel-automation\u002Freading-excel-files-with-pandas\u002F","Reading Excel Files with pandas",".",[29,30,39,40,39,44,39,48,39,55,39,61,39,66,39,71,39,80,39,85,39,89,39,92,39,95,39,99,39,102,39,105,39,109,39,112,39,115,39,120,39,126,39,129,39,133,39,138,39,142,39,150,39,155,39,160,39,163,39,166,39,169,39,171],"svg",{"viewBox":31,"role":32,"ariaLabel":33,"ariaLabelledBy":34,"xmlns":37,"style":38},"0 0 800 262","img","An Excel sheet with a three-row title block above the real header row, showing which rows skiprows discards and which row the header argument selects.",[35,36],"hdr-t","hdr-d","http:\u002F\u002Fwww.w3.org\u002F2000\u002Fsvg","width:100%;max-width:800px;height:auto;display:block;margin:1.5rem auto;font-family:Inter,ui-sans-serif,system-ui,sans-serif","\n  ",[41,42,43],"title",{"id":35},"Where the preamble ends and the header begins",[45,46,47],"desc",{"id":36},"A sheet layout with six rows. Rows zero through two hold a company banner, a report title and an as-at date. Row three is blank. Row four holds the real column names Region, Units and Revenue. Row five onwards holds data. Passing header equals four tells pandas to take row four as the column names and start data at row five, discarding everything above.",[49,50],"rect",{"x":51,"y":51,"width":52,"height":53,"fill":54},"0","800","262","#ffffff",[56,57,60],"text",{"x":58,"y":58,"style":59},"34","font-size:11px;font-weight:700;fill:var(--muted,#5b6780)","row",[56,62,65],{"x":63,"y":58,"style":64},"440","font-size:11px;font-weight:700;fill:var(--muted,#5b6780);text-anchor:middle","what pandas does with it",[56,67,51],{"x":68,"y":69,"style":70},"38","66","font-size:11px;fill:var(--muted,#5b6780);text-anchor:middle",[49,72],{"x":73,"y":74,"width":75,"height":76,"rx":77,"fill":78,"stroke":79},"58","48","290","26","5","#f0f2f5","var(--line,#cdd5e6)",[56,81,84],{"x":82,"y":69,"style":83},"203","font-size:11px;fill:var(--text,#172033);text-anchor:middle","ACME Corporation",[56,86,88],{"x":68,"y":87,"style":70},"98","1",[49,90],{"x":73,"y":91,"width":75,"height":76,"rx":77,"fill":78,"stroke":79},"80",[56,93,94],{"x":82,"y":87,"style":83},"Regional revenue report",[56,96,98],{"x":68,"y":97,"style":70},"130","2",[49,100],{"x":73,"y":101,"width":75,"height":76,"rx":77,"fill":78,"stroke":79},"112",[56,103,104],{"x":82,"y":97,"style":83},"As at 15 August 2026",[56,106,108],{"x":68,"y":107,"style":70},"162","3",[49,110],{"x":73,"y":111,"width":75,"height":76,"rx":77,"fill":78,"stroke":79},"144",[56,113,114],{"x":82,"y":107,"style":70},"(blank)",[56,116,119],{"x":68,"y":117,"style":118},"194","font-size:11px;font-weight:700;fill:var(--brand-strong,#4338ca);text-anchor:middle","4",[49,121],{"x":73,"y":122,"width":75,"height":76,"rx":77,"fill":123,"stroke":124,"style":125},"176","#ebebfd","var(--brand,#5b5cf0)","stroke-width:2px",[56,127,128],{"x":82,"y":117,"style":118},"Region · Units · Revenue",[56,130,132],{"x":68,"y":131,"style":70},"226","5+",[49,134],{"x":73,"y":135,"width":75,"height":76,"rx":77,"fill":136,"stroke":137,"style":125},"208","#d9f4f1","var(--teal,#0f9488)",[56,139,141],{"x":82,"y":131,"style":140},"font-size:11px;font-weight:700;fill:var(--teal-ink,#0b6157);text-anchor:middle","North · 412 · 5150.00",[49,143],{"x":144,"y":74,"width":145,"height":146,"rx":147,"fill":148,"stroke":149,"style":125},"380","404","122","12","#fee8f2","var(--accent,#f43f8f)",[56,151,154],{"x":152,"y":87,"style":153},"582","font-size:12px;font-weight:700;fill:var(--accent-ink,#be185d);text-anchor:middle","discarded — preamble",[56,156,159],{"x":152,"y":157,"style":158},"126","font-size:10.5px;fill:var(--muted,#5b6780);text-anchor:middle","read from row 0 and these become",[56,161,162],{"x":152,"y":111,"style":158},"your column names instead",[49,164],{"x":144,"y":122,"width":145,"height":76,"rx":165,"fill":123,"stroke":124,"style":125},"6",[56,167,168],{"x":152,"y":117,"style":118},"header=4 → the column names",[49,170],{"x":144,"y":135,"width":145,"height":76,"rx":165,"fill":136,"stroke":137,"style":125},[56,172,173],{"x":152,"y":131,"style":140},"data starts here automatically",[175,176,178],"h2",{"id":177},"prerequisites","Prerequisites",[180,181,186],"pre",{"className":182,"code":183,"language":184,"meta":185,"style":185},"language-bash shiki shiki-themes github-light github-dark-high-contrast","pip install pandas openpyxl\n","bash","",[14,187,188],{"__ignoreMap":185},[189,190,193,197,201,204],"span",{"class":191,"line":192},"line",1,[189,194,196],{"class":195},"sMTad","pip",[189,198,200],{"class":199},"srMev"," install",[189,202,203],{"class":199}," pandas",[189,205,206],{"class":199}," openpyxl\n",[10,208,209],{},"A file shaped like the ones that cause the problem:",[180,211,215],{"className":212,"code":213,"language":214,"meta":185,"style":185},"language-python shiki shiki-themes github-light github-dark-high-contrast","import pandas as pd\n\nwith pd.ExcelWriter(\"export.xlsx\", engine=\"xlsxwriter\") as writer:\n    frame = pd.DataFrame({\n        \"Region\": [\"North\", \"South\", \"West\"],\n        \"Units\": [412, 388, 265],\n        \"Revenue\": [5150.00, 4268.00, 3511.25],\n    })\n    frame.to_excel(writer, sheet_name=\"Report\", index=False, startrow=4)\n\n    sheet = writer.sheets[\"Report\"]\n    sheet.write(0, 0, \"ACME Corporation\")\n    sheet.write(1, 0, \"Regional revenue report\")\n    sheet.write(2, 0, \"As at 15 August 2026\")\n","python",[14,216,217,233,240,273,284,309,333,356,362,398,403,419,438,456],{"__ignoreMap":185},[189,218,219,223,227,230],{"class":191,"line":192},[189,220,222],{"class":221},"s-kum","import",[189,224,226],{"class":225},"skGVy"," pandas ",[189,228,229],{"class":221},"as",[189,231,232],{"class":225}," pd\n",[189,234,236],{"class":191,"line":235},2,[189,237,239],{"emptyLinePlaceholder":238},true,"\n",[189,241,243,246,249,252,255,259,262,265,268,270],{"class":191,"line":242},3,[189,244,245],{"class":221},"with",[189,247,248],{"class":225}," pd.ExcelWriter(",[189,250,251],{"class":199},"\"export.xlsx\"",[189,253,254],{"class":225},", ",[189,256,258],{"class":257},"sa561","engine",[189,260,261],{"class":221},"=",[189,263,264],{"class":199},"\"xlsxwriter\"",[189,266,267],{"class":225},") ",[189,269,229],{"class":221},[189,271,272],{"class":225}," writer:\n",[189,274,276,279,281],{"class":191,"line":275},4,[189,277,278],{"class":225},"    frame ",[189,280,261],{"class":221},[189,282,283],{"class":225}," pd.DataFrame({\n",[189,285,287,290,293,296,298,301,303,306],{"class":191,"line":286},5,[189,288,289],{"class":199},"        \"Region\"",[189,291,292],{"class":225},": [",[189,294,295],{"class":199},"\"North\"",[189,297,254],{"class":225},[189,299,300],{"class":199},"\"South\"",[189,302,254],{"class":225},[189,304,305],{"class":199},"\"West\"",[189,307,308],{"class":225},"],\n",[189,310,312,315,317,321,323,326,328,331],{"class":191,"line":311},6,[189,313,314],{"class":199},"        \"Units\"",[189,316,292],{"class":225},[189,318,320],{"class":319},"sP0c6","412",[189,322,254],{"class":225},[189,324,325],{"class":319},"388",[189,327,254],{"class":225},[189,329,330],{"class":319},"265",[189,332,308],{"class":225},[189,334,336,339,341,344,346,349,351,354],{"class":191,"line":335},7,[189,337,338],{"class":199},"        \"Revenue\"",[189,340,292],{"class":225},[189,342,343],{"class":319},"5150.00",[189,345,254],{"class":225},[189,347,348],{"class":319},"4268.00",[189,350,254],{"class":225},[189,352,353],{"class":319},"3511.25",[189,355,308],{"class":225},[189,357,359],{"class":191,"line":358},8,[189,360,361],{"class":225},"    })\n",[189,363,365,368,371,373,376,378,381,383,386,388,391,393,395],{"class":191,"line":364},9,[189,366,367],{"class":225},"    frame.to_excel(writer, ",[189,369,370],{"class":257},"sheet_name",[189,372,261],{"class":221},[189,374,375],{"class":199},"\"Report\"",[189,377,254],{"class":225},[189,379,380],{"class":257},"index",[189,382,261],{"class":221},[189,384,385],{"class":319},"False",[189,387,254],{"class":225},[189,389,390],{"class":257},"startrow",[189,392,261],{"class":221},[189,394,119],{"class":319},[189,396,397],{"class":225},")\n",[189,399,401],{"class":191,"line":400},10,[189,402,239],{"emptyLinePlaceholder":238},[189,404,406,409,411,414,416],{"class":191,"line":405},11,[189,407,408],{"class":225},"    sheet ",[189,410,261],{"class":221},[189,412,413],{"class":225}," writer.sheets[",[189,415,375],{"class":199},[189,417,418],{"class":225},"]\n",[189,420,422,425,427,429,431,433,436],{"class":191,"line":421},12,[189,423,424],{"class":225},"    sheet.write(",[189,426,51],{"class":319},[189,428,254],{"class":225},[189,430,51],{"class":319},[189,432,254],{"class":225},[189,434,435],{"class":199},"\"ACME Corporation\"",[189,437,397],{"class":225},[189,439,441,443,445,447,449,451,454],{"class":191,"line":440},13,[189,442,424],{"class":225},[189,444,88],{"class":319},[189,446,254],{"class":225},[189,448,51],{"class":319},[189,450,254],{"class":225},[189,452,453],{"class":199},"\"Regional revenue report\"",[189,455,397],{"class":225},[189,457,459,461,463,465,467,469,472],{"class":191,"line":458},14,[189,460,424],{"class":225},[189,462,98],{"class":319},[189,464,254],{"class":225},[189,466,51],{"class":319},[189,468,254],{"class":225},[189,470,471],{"class":199},"\"As at 15 August 2026\"",[189,473,397],{"class":225},[175,475,477],{"id":476},"step-1-look-before-you-parse","Step 1 — Look before you parse",[10,479,480],{},"Never guess the layout. Read the top of the sheet with no header at all and print it:",[180,482,484],{"className":212,"code":483,"language":214,"meta":185,"style":185},"import pandas as pd\n\npeek = pd.read_excel(\"export.xlsx\", header=None, nrows=8)\nprint(peek.to_string())\n#          0        1        2\n# 0  ACME Corporation  NaN  NaN\n# 1  Regional revenue report  NaN  NaN\n# 2  As at 15 August 2026  NaN  NaN\n# 3  NaN  NaN  NaN\n# 4  Region  Units  Revenue\n# 5  North  412  5150.0\n",[14,485,486,496,500,534,542,548,553,558,563,568,573],{"__ignoreMap":185},[189,487,488,490,492,494],{"class":191,"line":192},[189,489,222],{"class":221},[189,491,226],{"class":225},[189,493,229],{"class":221},[189,495,232],{"class":225},[189,497,498],{"class":191,"line":235},[189,499,239],{"emptyLinePlaceholder":238},[189,501,502,505,507,510,512,514,517,519,522,524,527,529,532],{"class":191,"line":242},[189,503,504],{"class":225},"peek ",[189,506,261],{"class":221},[189,508,509],{"class":225}," pd.read_excel(",[189,511,251],{"class":199},[189,513,254],{"class":225},[189,515,516],{"class":257},"header",[189,518,261],{"class":221},[189,520,521],{"class":319},"None",[189,523,254],{"class":225},[189,525,526],{"class":257},"nrows",[189,528,261],{"class":221},[189,530,531],{"class":319},"8",[189,533,397],{"class":225},[189,535,536,539],{"class":191,"line":275},[189,537,538],{"class":319},"print",[189,540,541],{"class":225},"(peek.to_string())\n",[189,543,544],{"class":191,"line":286},[189,545,547],{"class":546},"s-wDw","#          0        1        2\n",[189,549,550],{"class":191,"line":311},[189,551,552],{"class":546},"# 0  ACME Corporation  NaN  NaN\n",[189,554,555],{"class":191,"line":335},[189,556,557],{"class":546},"# 1  Regional revenue report  NaN  NaN\n",[189,559,560],{"class":191,"line":358},[189,561,562],{"class":546},"# 2  As at 15 August 2026  NaN  NaN\n",[189,564,565],{"class":191,"line":364},[189,566,567],{"class":546},"# 3  NaN  NaN  NaN\n",[189,569,570],{"class":191,"line":400},[189,571,572],{"class":546},"# 4  Region  Units  Revenue\n",[189,574,575],{"class":191,"line":405},[189,576,577],{"class":546},"# 5  North  412  5150.0\n",[10,579,580,583,584,587],{},[14,581,582],{},"header=None"," stops pandas promoting anything to column names, and ",[14,585,586],{},"nrows=8"," keeps it cheap on a large file. The header is clearly row 4.",[175,589,591],{"id":590},"step-2-set-the-header-row","Step 2 — Set the header row",[10,593,594],{},"With the index known, one argument does the job:",[180,596,598],{"className":212,"code":597,"language":214,"meta":185,"style":185},"df = pd.read_excel(\"export.xlsx\", header=4)\nprint(df.columns.tolist())     # ['Region', 'Units', 'Revenue']\nprint(len(df))                 # 3\n",[14,599,600,621,631],{"__ignoreMap":185},[189,601,602,605,607,609,611,613,615,617,619],{"class":191,"line":192},[189,603,604],{"class":225},"df ",[189,606,261],{"class":221},[189,608,509],{"class":225},[189,610,251],{"class":199},[189,612,254],{"class":225},[189,614,516],{"class":257},[189,616,261],{"class":221},[189,618,119],{"class":319},[189,620,397],{"class":225},[189,622,623,625,628],{"class":191,"line":235},[189,624,538],{"class":319},[189,626,627],{"class":225},"(df.columns.tolist())     ",[189,629,630],{"class":546},"# ['Region', 'Units', 'Revenue']\n",[189,632,633,635,638,641,644],{"class":191,"line":242},[189,634,538],{"class":319},[189,636,637],{"class":225},"(",[189,639,640],{"class":319},"len",[189,642,643],{"class":225},"(df))                 ",[189,645,646],{"class":546},"# 3\n",[10,648,649,650,654,655,658,659,662,663,665,666,670],{},"Notice what you did ",[651,652,653],"em",{},"not"," need: ",[14,656,657],{},"skiprows",". Setting ",[14,660,661],{},"header=4"," already tells pandas to ignore rows 0 through 3 and start data at row 5. Combining both is the most common source of confusion, because ",[14,664,516],{}," is interpreted ",[667,668,669],"strong",{},"relative to what remains after skipping",":",[180,672,674],{"className":212,"code":673,"language":214,"meta":185,"style":185},"# Equivalent to header=4 — the counting restarts after the skip.\ndf = pd.read_excel(\"export.xlsx\", skiprows=4, header=0)\n\n# NOT equivalent: skips 4 rows, then takes the 5th remaining row as header,\n# which is the first data row. Almost never what you want.\ndf = pd.read_excel(\"export.xlsx\", skiprows=4, header=4)\n",[14,675,676,681,709,713,718,723],{"__ignoreMap":185},[189,677,678],{"class":191,"line":192},[189,679,680],{"class":546},"# Equivalent to header=4 — the counting restarts after the skip.\n",[189,682,683,685,687,689,691,693,695,697,699,701,703,705,707],{"class":191,"line":235},[189,684,604],{"class":225},[189,686,261],{"class":221},[189,688,509],{"class":225},[189,690,251],{"class":199},[189,692,254],{"class":225},[189,694,657],{"class":257},[189,696,261],{"class":221},[189,698,119],{"class":319},[189,700,254],{"class":225},[189,702,516],{"class":257},[189,704,261],{"class":221},[189,706,51],{"class":319},[189,708,397],{"class":225},[189,710,711],{"class":191,"line":242},[189,712,239],{"emptyLinePlaceholder":238},[189,714,715],{"class":191,"line":275},[189,716,717],{"class":546},"# NOT equivalent: skips 4 rows, then takes the 5th remaining row as header,\n",[189,719,720],{"class":191,"line":286},[189,721,722],{"class":546},"# which is the first data row. Almost never what you want.\n",[189,724,725,727,729,731,733,735,737,739,741,743,745,747,749],{"class":191,"line":311},[189,726,604],{"class":225},[189,728,261],{"class":221},[189,730,509],{"class":225},[189,732,251],{"class":199},[189,734,254],{"class":225},[189,736,657],{"class":257},[189,738,261],{"class":221},[189,740,119],{"class":319},[189,742,254],{"class":225},[189,744,516],{"class":257},[189,746,261],{"class":221},[189,748,119],{"class":319},[189,750,397],{"class":225},[10,752,753,754,761,762,764,765,767],{},"The rule to remember: ",[667,755,756,757,760],{},"use ",[14,758,759],{},"header="," alone when the preamble is simply above the header",". Reach for ",[14,763,657],{}," only when you need to discard rows that are ",[651,766,653],{}," contiguous with the top, which the callable form handles.",[769,770,771,784],"table",{},[772,773,774],"thead",{},[775,776,777,781],"tr",{},[778,779,780],"th",{},"Goal",[778,782,783],{},"Argument",[785,786,787,797,807,817,827,837],"tbody",{},[775,788,789,793],{},[790,791,792],"td",{},"Header is on row 4, preamble above",[790,794,795],{},[14,796,661],{},[775,798,799,802],{},[790,800,801],{},"Two stacked header rows",[790,803,804],{},[14,805,806],{},"header=[0, 1]",[775,808,809,812],{},[790,810,811],{},"No header at all; supply names",[790,813,814],{},[14,815,816],{},"header=None, names=[...]",[775,818,819,822],{},[790,820,821],{},"Drop scattered rows anywhere",[790,823,824],{},[14,825,826],{},"skiprows=lambda i: ...",[775,828,829,832],{},[790,830,831],{},"Drop trailing total rows",[790,833,834],{},[14,835,836],{},"skipfooter=2",[775,838,839,842],{},[790,840,841],{},"Read only the first 500 data rows",[790,843,844],{},[14,845,846],{},"nrows=500",[175,848,850],{"id":849},"step-3-handle-stacked-headers","Step 3 — Handle stacked headers",[29,852,39,858,39,861,39,864,39,866,39,872,39,878,39,884,39,886,39,891,39,895,39,900,39,906,39,909,39,912,39,915,39,919,39,921,39,924,39,927,39,929,39,932,39,935,39,940,39,944,39,946,39,949,39,951,39,953,39,961],{"viewBox":853,"role":32,"ariaLabel":854,"ariaLabelledBy":855,"xmlns":37,"style":38},"0 0 800 226","A two-row header where a merged group name spans three detail columns, producing one real name and two Unnamed placeholders that must be filled before flattening.",[856,857],"stack-t","stack-d",[41,859,860],{"id":856},"How a merged group header becomes Unnamed columns",[45,862,863],{"id":857},"The top header row holds Q1 spanning three columns as a merged cell, so only the leftmost of the three carries the text and the other two are blank. Read as a MultiIndex, those blanks become Unnamed placeholders. Forward-filling the upper level spreads Q1 rightwards, after which joining the two levels produces Q1_Units, Q1_Revenue and Q1_Margin.",[49,865],{"x":51,"y":51,"width":52,"height":131,"fill":54},[56,867,871],{"x":868,"y":869,"style":870},"400","28","font-size:12px;font-weight:700;fill:var(--muted,#5b6780);text-anchor:middle","row 0: the merged group header",[49,873],{"x":874,"y":875,"width":876,"height":877,"rx":165,"fill":123,"stroke":124,"style":125},"60","40","330","32",[56,879,883],{"x":880,"y":881,"style":882},"225","61","font-size:11.5px;font-weight:700;fill:var(--brand-strong,#4338ca);text-anchor:middle","Q1 — merged across three columns",[49,885],{"x":868,"y":875,"width":876,"height":877,"rx":165,"fill":136,"stroke":137,"style":125},[56,887,890],{"x":888,"y":881,"style":889},"565","font-size:11.5px;font-weight:700;fill:var(--teal-ink,#0b6157);text-anchor:middle","Q2 — merged across three columns",[56,892,894],{"x":868,"y":893,"style":870},"96","row 1: the detail header",[49,896],{"x":874,"y":897,"width":898,"height":899,"rx":77,"fill":78,"stroke":79},"108","106","30",[56,901,905],{"x":902,"y":903,"style":904},"113","128","font-size:10.5px;fill:var(--text,#172033);text-anchor:middle","Units",[49,907],{"x":908,"y":897,"width":898,"height":899,"rx":77,"fill":78,"stroke":79},"172",[56,910,911],{"x":880,"y":903,"style":904},"Revenue",[49,913],{"x":914,"y":897,"width":898,"height":899,"rx":77,"fill":78,"stroke":79},"284",[56,916,918],{"x":917,"y":903,"style":904},"337","Margin",[49,920],{"x":868,"y":897,"width":898,"height":899,"rx":77,"fill":78,"stroke":79},[56,922,905],{"x":923,"y":903,"style":904},"453",[49,925],{"x":926,"y":897,"width":898,"height":899,"rx":77,"fill":78,"stroke":79},"512",[56,928,911],{"x":888,"y":903,"style":904},[49,930],{"x":931,"y":897,"width":898,"height":899,"rx":77,"fill":78,"stroke":79},"624",[56,933,918],{"x":934,"y":903,"style":904},"677",[56,936,939],{"x":902,"y":937,"style":938},"160","font-size:10px;fill:var(--muted,#5b6780);text-anchor:middle","Q1",[56,941,943],{"x":880,"y":937,"style":942},"font-size:10px;fill:var(--accent-ink,#be185d);text-anchor:middle","Unnamed",[56,945,943],{"x":917,"y":937,"style":942},[56,947,948],{"x":923,"y":937,"style":938},"Q2",[56,950,943],{"x":888,"y":937,"style":942},[56,952,943],{"x":934,"y":937,"style":942},[49,954],{"x":874,"y":955,"width":956,"height":957,"rx":958,"fill":959,"stroke":960,"style":125},"174","670","36","9","#fdefd8","var(--gold,#b4740a)",[56,962,966],{"x":963,"y":964,"style":965},"395","197","font-size:11px;font-weight:700;fill:var(--gold-ink,#7a4e06);text-anchor:middle","fill the upper level rightwards, then join → Q1_Units, Q1_Revenue, Q1_Margin, Q2_Units …",[10,968,969,970,972,973,254,975,254,977,979,980,670],{},"Exports from reporting tools often stack a group row above a detail row — ",[14,971,939],{}," spanning three columns, then ",[14,974,905],{},[14,976,911],{},[14,978,918],{}," beneath. Pass a list and pandas builds a ",[14,981,982],{},"MultiIndex",[180,984,986],{"className":212,"code":985,"language":214,"meta":185,"style":185},"import pandas as pd\n\ndf = pd.read_excel(\"quarterly.xlsx\", header=[0, 1])\nprint(df.columns[:3].tolist())\n# [('Q1', 'Units'), ('Q1', 'Revenue'), ('Q1', 'Margin')]\n",[14,987,988,998,1002,1031,1043],{"__ignoreMap":185},[189,989,990,992,994,996],{"class":191,"line":192},[189,991,222],{"class":221},[189,993,226],{"class":225},[189,995,229],{"class":221},[189,997,232],{"class":225},[189,999,1000],{"class":191,"line":235},[189,1001,239],{"emptyLinePlaceholder":238},[189,1003,1004,1006,1008,1010,1013,1015,1017,1019,1022,1024,1026,1028],{"class":191,"line":242},[189,1005,604],{"class":225},[189,1007,261],{"class":221},[189,1009,509],{"class":225},[189,1011,1012],{"class":199},"\"quarterly.xlsx\"",[189,1014,254],{"class":225},[189,1016,516],{"class":257},[189,1018,261],{"class":221},[189,1020,1021],{"class":225},"[",[189,1023,51],{"class":319},[189,1025,254],{"class":225},[189,1027,88],{"class":319},[189,1029,1030],{"class":225},"])\n",[189,1032,1033,1035,1038,1040],{"class":191,"line":275},[189,1034,538],{"class":319},[189,1036,1037],{"class":225},"(df.columns[:",[189,1039,108],{"class":319},[189,1041,1042],{"class":225},"].tolist())\n",[189,1044,1045],{"class":191,"line":286},[189,1046,1047],{"class":546},"# [('Q1', 'Units'), ('Q1', 'Revenue'), ('Q1', 'Margin')]\n",[10,1049,1050,1051,1053,1054,1057],{},"A ",[14,1052,982],{}," is awkward to work with downstream, so flatten it immediately. Merged group cells leave ",[14,1055,1056],{},"Unnamed:"," fragments in the upper level, which need dropping as you join:",[180,1059,1061],{"className":212,"code":1060,"language":214,"meta":185,"style":185},"def flatten(columns):\n    \"\"\"Join MultiIndex levels, ignoring the Unnamed fragments merges leave.\"\"\"\n    flat = []\n    for parts in columns:\n        keep = [\n            str(p).strip() for p in parts\n            if p is not None and not str(p).startswith(\"Unnamed:\")\n        ]\n        flat.append(\"_\".join(keep) if keep else \"unnamed\")\n    return flat\n\ndf.columns = flatten(df.columns)\nprint(df.columns.tolist())\n# ['Q1_Units', 'Q1_Revenue', 'Q1_Margin', 'Q2_Units', ...]\n",[14,1062,1063,1075,1080,1090,1104,1114,1133,1165,1170,1195,1203,1207,1217,1224],{"__ignoreMap":185},[189,1064,1065,1068,1072],{"class":191,"line":192},[189,1066,1067],{"class":221},"def",[189,1069,1071],{"class":1070},"s_Opv"," flatten",[189,1073,1074],{"class":225},"(columns):\n",[189,1076,1077],{"class":191,"line":235},[189,1078,1079],{"class":199},"    \"\"\"Join MultiIndex levels, ignoring the Unnamed fragments merges leave.\"\"\"\n",[189,1081,1082,1085,1087],{"class":191,"line":242},[189,1083,1084],{"class":225},"    flat ",[189,1086,261],{"class":221},[189,1088,1089],{"class":225}," []\n",[189,1091,1092,1095,1098,1101],{"class":191,"line":275},[189,1093,1094],{"class":221},"    for",[189,1096,1097],{"class":225}," parts ",[189,1099,1100],{"class":221},"in",[189,1102,1103],{"class":225}," columns:\n",[189,1105,1106,1109,1111],{"class":191,"line":286},[189,1107,1108],{"class":225},"        keep ",[189,1110,261],{"class":221},[189,1112,1113],{"class":225}," [\n",[189,1115,1116,1119,1122,1125,1128,1130],{"class":191,"line":311},[189,1117,1118],{"class":319},"            str",[189,1120,1121],{"class":225},"(p).strip() ",[189,1123,1124],{"class":221},"for",[189,1126,1127],{"class":225}," p ",[189,1129,1100],{"class":221},[189,1131,1132],{"class":225}," parts\n",[189,1134,1135,1138,1140,1143,1146,1149,1152,1154,1157,1160,1163],{"class":191,"line":335},[189,1136,1137],{"class":221},"            if",[189,1139,1127],{"class":225},[189,1141,1142],{"class":221},"is",[189,1144,1145],{"class":221}," not",[189,1147,1148],{"class":319}," None",[189,1150,1151],{"class":221}," and",[189,1153,1145],{"class":221},[189,1155,1156],{"class":319}," str",[189,1158,1159],{"class":225},"(p).startswith(",[189,1161,1162],{"class":199},"\"Unnamed:\"",[189,1164,397],{"class":225},[189,1166,1167],{"class":191,"line":358},[189,1168,1169],{"class":225},"        ]\n",[189,1171,1172,1175,1178,1181,1184,1187,1190,1193],{"class":191,"line":364},[189,1173,1174],{"class":225},"        flat.append(",[189,1176,1177],{"class":199},"\"_\"",[189,1179,1180],{"class":225},".join(keep) ",[189,1182,1183],{"class":221},"if",[189,1185,1186],{"class":225}," keep ",[189,1188,1189],{"class":221},"else",[189,1191,1192],{"class":199}," \"unnamed\"",[189,1194,397],{"class":225},[189,1196,1197,1200],{"class":191,"line":400},[189,1198,1199],{"class":221},"    return",[189,1201,1202],{"class":225}," flat\n",[189,1204,1205],{"class":191,"line":405},[189,1206,239],{"emptyLinePlaceholder":238},[189,1208,1209,1212,1214],{"class":191,"line":421},[189,1210,1211],{"class":225},"df.columns ",[189,1213,261],{"class":221},[189,1215,1216],{"class":225}," flatten(df.columns)\n",[189,1218,1219,1221],{"class":191,"line":440},[189,1220,538],{"class":319},[189,1222,1223],{"class":225},"(df.columns.tolist())\n",[189,1225,1226],{"class":191,"line":458},[189,1227,1228],{"class":546},"# ['Q1_Units', 'Q1_Revenue', 'Q1_Margin', 'Q2_Units', ...]\n",[10,1230,1231,1232,1234,1235,1237,1238,27],{},"The ",[14,1233,1056],{}," filter is essential because Excel stores a merged cell's value only in its top-left cell — the rest read as blank, so a header spanning three columns produces one real name and two ",[14,1236,1056],{}," placeholders. The wider treatment of that behaviour is in ",[23,1239,1241],{"href":1240},"\u002Fgetting-started-with-python-excel-automation\u002Freading-excel-files-with-pandas\u002Fhandle-merged-cells-when-reading-excel-with-pandas\u002F","handling merged cells when reading Excel with pandas",[175,1243,1245],{"id":1244},"step-4-find-the-header-row-automatically","Step 4 — Find the header row automatically",[10,1247,1248,1249,1251],{},"Hard-coding ",[14,1250,661],{}," works until the month somebody adds a line to the title block. The durable answer is to search for the header by its content:",[29,1253,39,1260,39,1263,39,1266,39,1271,39,1277,39,1281,39,1283,39,1286,39,1289,39,1294,39,1298,39,1304,39,1308,39,1312,39,1316,39,1320,39,1324,39,1328,39,1331,39,1334],{"viewBox":1254,"role":32,"ariaLabel":1255,"ariaLabelledBy":1256,"xmlns":37,"style":1259},"-2 34 802 181","Header detection: scan the first rows for the one containing the known column names, then re-read the file passing that index as header.",[1257,1258],"find-t","find-d","width:100%;max-width:802px;height:auto;display:block;margin:1.5rem auto;font-family:Inter,ui-sans-serif,system-ui,sans-serif",[41,1261,1262],{"id":1257},"Detecting the header row instead of hard-coding it",[45,1264,1265],{"id":1258},"A two-pass read. The first pass reads the top twenty rows with header set to None, and scans each row's cell values for a required set of known column names such as Region and Revenue. The index of the first matching row becomes the header argument for a second, full read. This survives a title block that gains or loses a line between months, which a hard-coded index does not.",[49,1267],{"x":1268,"y":58,"width":1269,"height":1270,"fill":54},"-2","802","181",[49,1272],{"x":1273,"y":1274,"width":1275,"height":1274,"rx":1276,"fill":123,"stroke":124,"style":125},"14","76","180","13",[56,1278,1280],{"x":1279,"y":874,"style":64},"104","pass 1",[56,1282,582],{"x":1279,"y":898,"style":882},[56,1284,1285],{"x":1279,"y":903,"style":158},"nrows=20 — cheap",[191,1287],{"x1":117,"y1":1288,"x2":131,"y2":1288,"stroke":124,"style":125},"114",[1290,1291],"polygon",{"points":1292,"fill":1293},"234,114 222,108 222,120","#5b5cf0",[49,1295],{"x":1296,"y":69,"width":1297,"height":893,"rx":1276,"fill":959,"stroke":960,"style":125},"242","220",[56,1299,1303],{"x":1300,"y":1301,"style":1302},"352","94","font-size:11.5px;font-weight:700;fill:var(--gold-ink,#7a4e06);text-anchor:middle","scan each row",[56,1305,1307],{"x":1300,"y":1306,"style":904},"118","does it contain",[56,1309,1311],{"x":1300,"y":1310,"style":904},"136","Region and Revenue?",[191,1313],{"x1":1314,"y1":1288,"x2":1315,"y2":1288,"stroke":960,"style":125},"462","494",[1290,1317],{"points":1318,"fill":1319},"502,114 490,108 490,120","#b4740a",[49,1321],{"x":1322,"y":69,"width":1323,"height":893,"rx":1276,"fill":136,"stroke":137,"style":125},"510","274",[56,1325,1327],{"x":1326,"y":1301,"style":889},"647","pass 2: header=\u003Cthat index>",[56,1329,1330],{"x":1326,"y":1306,"style":904},"survives a title block that",[56,1332,1333],{"x":1326,"y":1310,"style":904},"gains or loses a line",[56,1335,1337],{"x":868,"y":1336,"style":70},"196","raise a clear error when no row matches — a silent fallback hides a changed export",[180,1339,1341],{"className":212,"code":1340,"language":214,"meta":185,"style":185},"import pandas as pd\n\ndef find_header_row(path, required, sheet_name=0, search=20):\n    \"\"\"Return the index of the first row containing all required column names.\"\"\"\n    wanted = {str(name).strip().lower() for name in required}\n\n    preview = pd.read_excel(path, sheet_name=sheet_name,\n                            header=None, nrows=search)\n\n    for index, row in preview.iterrows():\n        values = {str(v).strip().lower() for v in row if pd.notna(v)}\n        if wanted \u003C= values:\n            return int(index)\n\n    raise ValueError(\n        f\"No header row in the first {search} rows of {path} contains \"\n        f\"{sorted(required)} — has the export format changed?\"\n    )\n\ndef read_report(path, required, sheet_name=0, **kwargs):\n    header = find_header_row(path, required, sheet_name=sheet_name)\n    return pd.read_excel(path, sheet_name=sheet_name, header=header, **kwargs)\n\ndf = read_report(\"export.xlsx\", [\"Region\", \"Units\", \"Revenue\"])\nprint(df.head())\n",[14,1342,1343,1353,1357,1382,1387,1413,1417,1434,1452,1456,1468,1497,1511,1522,1526,1538,1570,1591,1597,1602,1624,1642,1668,1673,1703],{"__ignoreMap":185},[189,1344,1345,1347,1349,1351],{"class":191,"line":192},[189,1346,222],{"class":221},[189,1348,226],{"class":225},[189,1350,229],{"class":221},[189,1352,232],{"class":225},[189,1354,1355],{"class":191,"line":235},[189,1356,239],{"emptyLinePlaceholder":238},[189,1358,1359,1361,1364,1367,1369,1371,1374,1376,1379],{"class":191,"line":242},[189,1360,1067],{"class":221},[189,1362,1363],{"class":1070}," find_header_row",[189,1365,1366],{"class":225},"(path, required, sheet_name",[189,1368,261],{"class":221},[189,1370,51],{"class":319},[189,1372,1373],{"class":225},", search",[189,1375,261],{"class":221},[189,1377,1378],{"class":319},"20",[189,1380,1381],{"class":225},"):\n",[189,1383,1384],{"class":191,"line":275},[189,1385,1386],{"class":199},"    \"\"\"Return the index of the first row containing all required column names.\"\"\"\n",[189,1388,1389,1392,1394,1397,1400,1403,1405,1408,1410],{"class":191,"line":286},[189,1390,1391],{"class":225},"    wanted ",[189,1393,261],{"class":221},[189,1395,1396],{"class":225}," {",[189,1398,1399],{"class":319},"str",[189,1401,1402],{"class":225},"(name).strip().lower() ",[189,1404,1124],{"class":221},[189,1406,1407],{"class":225}," name ",[189,1409,1100],{"class":221},[189,1411,1412],{"class":225}," required}\n",[189,1414,1415],{"class":191,"line":311},[189,1416,239],{"emptyLinePlaceholder":238},[189,1418,1419,1422,1424,1427,1429,1431],{"class":191,"line":335},[189,1420,1421],{"class":225},"    preview ",[189,1423,261],{"class":221},[189,1425,1426],{"class":225}," pd.read_excel(path, ",[189,1428,370],{"class":257},[189,1430,261],{"class":221},[189,1432,1433],{"class":225},"sheet_name,\n",[189,1435,1436,1439,1441,1443,1445,1447,1449],{"class":191,"line":358},[189,1437,1438],{"class":257},"                            header",[189,1440,261],{"class":221},[189,1442,521],{"class":319},[189,1444,254],{"class":225},[189,1446,526],{"class":257},[189,1448,261],{"class":221},[189,1450,1451],{"class":225},"search)\n",[189,1453,1454],{"class":191,"line":364},[189,1455,239],{"emptyLinePlaceholder":238},[189,1457,1458,1460,1463,1465],{"class":191,"line":400},[189,1459,1094],{"class":221},[189,1461,1462],{"class":225}," index, row ",[189,1464,1100],{"class":221},[189,1466,1467],{"class":225}," preview.iterrows():\n",[189,1469,1470,1473,1475,1477,1479,1482,1484,1487,1489,1492,1494],{"class":191,"line":405},[189,1471,1472],{"class":225},"        values ",[189,1474,261],{"class":221},[189,1476,1396],{"class":225},[189,1478,1399],{"class":319},[189,1480,1481],{"class":225},"(v).strip().lower() ",[189,1483,1124],{"class":221},[189,1485,1486],{"class":225}," v ",[189,1488,1100],{"class":221},[189,1490,1491],{"class":225}," row ",[189,1493,1183],{"class":221},[189,1495,1496],{"class":225}," pd.notna(v)}\n",[189,1498,1499,1502,1505,1508],{"class":191,"line":421},[189,1500,1501],{"class":221},"        if",[189,1503,1504],{"class":225}," wanted ",[189,1506,1507],{"class":221},"\u003C=",[189,1509,1510],{"class":225}," values:\n",[189,1512,1513,1516,1519],{"class":191,"line":440},[189,1514,1515],{"class":221},"            return",[189,1517,1518],{"class":319}," int",[189,1520,1521],{"class":225},"(index)\n",[189,1523,1524],{"class":191,"line":458},[189,1525,239],{"emptyLinePlaceholder":238},[189,1527,1529,1532,1535],{"class":191,"line":1528},15,[189,1530,1531],{"class":221},"    raise",[189,1533,1534],{"class":319}," ValueError",[189,1536,1537],{"class":225},"(\n",[189,1539,1541,1544,1547,1551,1554,1557,1560,1562,1565,1567],{"class":191,"line":1540},16,[189,1542,1543],{"class":221},"        f",[189,1545,1546],{"class":199},"\"No header row in the first ",[189,1548,1550],{"class":1549},"sSjpA","{",[189,1552,1553],{"class":225},"search",[189,1555,1556],{"class":1549},"}",[189,1558,1559],{"class":199}," rows of ",[189,1561,1550],{"class":1549},[189,1563,1564],{"class":225},"path",[189,1566,1556],{"class":1549},[189,1568,1569],{"class":199}," contains \"\n",[189,1571,1573,1575,1578,1580,1583,1586,1588],{"class":191,"line":1572},17,[189,1574,1543],{"class":221},[189,1576,1577],{"class":199},"\"",[189,1579,1550],{"class":1549},[189,1581,1582],{"class":319},"sorted",[189,1584,1585],{"class":225},"(required)",[189,1587,1556],{"class":1549},[189,1589,1590],{"class":199}," — has the export format changed?\"\n",[189,1592,1594],{"class":191,"line":1593},18,[189,1595,1596],{"class":225},"    )\n",[189,1598,1600],{"class":191,"line":1599},19,[189,1601,239],{"emptyLinePlaceholder":238},[189,1603,1605,1607,1610,1612,1614,1616,1618,1621],{"class":191,"line":1604},20,[189,1606,1067],{"class":221},[189,1608,1609],{"class":1070}," read_report",[189,1611,1366],{"class":225},[189,1613,261],{"class":221},[189,1615,51],{"class":319},[189,1617,254],{"class":225},[189,1619,1620],{"class":221},"**",[189,1622,1623],{"class":225},"kwargs):\n",[189,1625,1627,1630,1632,1635,1637,1639],{"class":191,"line":1626},21,[189,1628,1629],{"class":225},"    header ",[189,1631,261],{"class":221},[189,1633,1634],{"class":225}," find_header_row(path, required, ",[189,1636,370],{"class":257},[189,1638,261],{"class":221},[189,1640,1641],{"class":225},"sheet_name)\n",[189,1643,1645,1647,1649,1651,1653,1656,1658,1660,1663,1665],{"class":191,"line":1644},22,[189,1646,1199],{"class":221},[189,1648,1426],{"class":225},[189,1650,370],{"class":257},[189,1652,261],{"class":221},[189,1654,1655],{"class":225},"sheet_name, ",[189,1657,516],{"class":257},[189,1659,261],{"class":221},[189,1661,1662],{"class":225},"header, ",[189,1664,1620],{"class":221},[189,1666,1667],{"class":225},"kwargs)\n",[189,1669,1671],{"class":191,"line":1670},23,[189,1672,239],{"emptyLinePlaceholder":238},[189,1674,1676,1678,1680,1683,1685,1688,1691,1693,1696,1698,1701],{"class":191,"line":1675},24,[189,1677,604],{"class":225},[189,1679,261],{"class":221},[189,1681,1682],{"class":225}," read_report(",[189,1684,251],{"class":199},[189,1686,1687],{"class":225},", [",[189,1689,1690],{"class":199},"\"Region\"",[189,1692,254],{"class":225},[189,1694,1695],{"class":199},"\"Units\"",[189,1697,254],{"class":225},[189,1699,1700],{"class":199},"\"Revenue\"",[189,1702,1030],{"class":225},[189,1704,1706,1708],{"class":191,"line":1705},25,[189,1707,538],{"class":319},[189,1709,1710],{"class":225},"(df.head())\n",[10,1712,1713,1714,1717,1718,1720,1721,27],{},"Raising when nothing matches is deliberate. A silent fallback to ",[14,1715,1716],{},"header=0"," produces a frame full of ",[14,1719,1056],{}," columns that fails confusingly three steps later, whereas this error names the file and the columns it expected. That is the same principle behind ",[23,1722,1724],{"href":1723},"\u002Fadvanced-data-transformation-and-cleaning\u002Fvalidating-excel-data-with-python\u002Fvalidate-excel-columns-before-import-with-pandas\u002F","validating Excel columns before import",[175,1726,1728],{"id":1727},"step-5-drop-trailing-rows-and-pick-columns","Step 5 — Drop trailing rows and pick columns",[10,1730,1731,1732,1735],{},"Exports often end with a blank line and a grand total. ",[14,1733,1734],{},"skipfooter"," removes them:",[180,1737,1739],{"className":212,"code":1738,"language":214,"meta":185,"style":185},"df = pd.read_excel(\"export.xlsx\", header=4, skipfooter=2)\n",[14,1740,1741],{"__ignoreMap":185},[189,1742,1743,1745,1747,1749,1751,1753,1755,1757,1759,1761,1763,1765,1767],{"class":191,"line":192},[189,1744,604],{"class":225},[189,1746,261],{"class":221},[189,1748,509],{"class":225},[189,1750,251],{"class":199},[189,1752,254],{"class":225},[189,1754,516],{"class":257},[189,1756,261],{"class":221},[189,1758,119],{"class":319},[189,1760,254],{"class":225},[189,1762,1734],{"class":257},[189,1764,261],{"class":221},[189,1766,98],{"class":319},[189,1768,397],{"class":225},[10,1770,1771,1772,1774],{},"Be aware of the cost: ",[14,1773,1734],{}," forces pandas down a slower Python-level path, because it cannot know where the end is until it has read everything. On a large sheet it is faster to read normally and slice:",[180,1776,1778],{"className":212,"code":1777,"language":214,"meta":185,"style":185},"df = pd.read_excel(\"export.xlsx\", header=4)\ndf = df.iloc[:-2]                                   # drop the last two rows\n\n# Better still, drop by content rather than position.\ndf = df[df[\"Region\"].notna() & (df[\"Region\"] != \"Total\")]\n",[14,1779,1780,1800,1820,1824,1829],{"__ignoreMap":185},[189,1781,1782,1784,1786,1788,1790,1792,1794,1796,1798],{"class":191,"line":192},[189,1783,604],{"class":225},[189,1785,261],{"class":221},[189,1787,509],{"class":225},[189,1789,251],{"class":199},[189,1791,254],{"class":225},[189,1793,516],{"class":257},[189,1795,261],{"class":221},[189,1797,119],{"class":319},[189,1799,397],{"class":225},[189,1801,1802,1804,1806,1809,1812,1814,1817],{"class":191,"line":235},[189,1803,604],{"class":225},[189,1805,261],{"class":221},[189,1807,1808],{"class":225}," df.iloc[:",[189,1810,1811],{"class":221},"-",[189,1813,98],{"class":319},[189,1815,1816],{"class":225},"]                                   ",[189,1818,1819],{"class":546},"# drop the last two rows\n",[189,1821,1822],{"class":191,"line":242},[189,1823,239],{"emptyLinePlaceholder":238},[189,1825,1826],{"class":191,"line":275},[189,1827,1828],{"class":546},"# Better still, drop by content rather than position.\n",[189,1830,1831,1833,1835,1838,1840,1843,1846,1849,1851,1854,1857,1860],{"class":191,"line":286},[189,1832,604],{"class":225},[189,1834,261],{"class":221},[189,1836,1837],{"class":225}," df[df[",[189,1839,1690],{"class":199},[189,1841,1842],{"class":225},"].notna() ",[189,1844,1845],{"class":221},"&",[189,1847,1848],{"class":225}," (df[",[189,1850,1690],{"class":199},[189,1852,1853],{"class":225},"] ",[189,1855,1856],{"class":221},"!=",[189,1858,1859],{"class":199}," \"Total\"",[189,1861,1862],{"class":225},")]\n",[10,1864,1865],{},"Dropping by content survives a month where the export has one trailing row instead of two, which position-based slicing does not.",[10,1867,1868],{},"Finally, read only the columns you need. It is faster and it removes a whole class of surprise from columns you never look at:",[180,1870,1872],{"className":212,"code":1871,"language":214,"meta":185,"style":185},"df = pd.read_excel(\"export.xlsx\", header=4, usecols=[\"Region\", \"Revenue\"])\ndf = pd.read_excel(\"export.xlsx\", header=4, usecols=\"A:C\")      # by letter\ndf = pd.read_excel(\"export.xlsx\", header=4, usecols=lambda c: not c.startswith(\"_\"))\n",[14,1873,1874,1909,1942],{"__ignoreMap":185},[189,1875,1876,1878,1880,1882,1884,1886,1888,1890,1892,1894,1897,1899,1901,1903,1905,1907],{"class":191,"line":192},[189,1877,604],{"class":225},[189,1879,261],{"class":221},[189,1881,509],{"class":225},[189,1883,251],{"class":199},[189,1885,254],{"class":225},[189,1887,516],{"class":257},[189,1889,261],{"class":221},[189,1891,119],{"class":319},[189,1893,254],{"class":225},[189,1895,1896],{"class":257},"usecols",[189,1898,261],{"class":221},[189,1900,1021],{"class":225},[189,1902,1690],{"class":199},[189,1904,254],{"class":225},[189,1906,1700],{"class":199},[189,1908,1030],{"class":225},[189,1910,1911,1913,1915,1917,1919,1921,1923,1925,1927,1929,1931,1933,1936,1939],{"class":191,"line":235},[189,1912,604],{"class":225},[189,1914,261],{"class":221},[189,1916,509],{"class":225},[189,1918,251],{"class":199},[189,1920,254],{"class":225},[189,1922,516],{"class":257},[189,1924,261],{"class":221},[189,1926,119],{"class":319},[189,1928,254],{"class":225},[189,1930,1896],{"class":257},[189,1932,261],{"class":221},[189,1934,1935],{"class":199},"\"A:C\"",[189,1937,1938],{"class":225},")      ",[189,1940,1941],{"class":546},"# by letter\n",[189,1943,1944,1946,1948,1950,1952,1954,1956,1958,1960,1962,1964,1967,1970,1972,1975,1977],{"class":191,"line":242},[189,1945,604],{"class":225},[189,1947,261],{"class":221},[189,1949,509],{"class":225},[189,1951,251],{"class":199},[189,1953,254],{"class":225},[189,1955,516],{"class":257},[189,1957,261],{"class":221},[189,1959,119],{"class":319},[189,1961,254],{"class":225},[189,1963,1896],{"class":257},[189,1965,1966],{"class":221},"=lambda",[189,1968,1969],{"class":225}," c: ",[189,1971,653],{"class":221},[189,1973,1974],{"class":225}," c.startswith(",[189,1976,1177],{"class":199},[189,1978,1979],{"class":225},"))\n",[175,1981,1983],{"id":1982},"common-pitfalls-and-fixes","Common pitfalls and fixes",[769,1985,1986,1999],{},[772,1987,1988],{},[775,1989,1990,1993,1996],{},[778,1991,1992],{},"Symptom",[778,1994,1995],{},"Cause",[778,1997,1998],{},"Fix",[785,2000,2001,2020,2042,2058,2078,2089,2102,2118],{},[775,2002,2003,2011,2014],{},[790,2004,2005,2006,254,2008],{},"Columns named ",[14,2007,16],{},[14,2009,2010],{},"Unnamed: 1",[790,2012,2013],{},"Header index points at a blank row",[790,2015,2016,2017,2019],{},"Peek with ",[14,2018,582],{}," and set the right index.",[775,2021,2022,2025,2033],{},[790,2023,2024],{},"Data missing its first row",[790,2026,2027,2029,2030,2032],{},[14,2028,657],{}," and ",[14,2031,516],{}," both set",[790,2034,2035,2036,2038,2039,2041],{},"Use ",[14,2037,759],{}," alone; ",[14,2040,516],{}," counts after the skip.",[775,2043,2044,2047,2052],{},[790,2045,2046],{},"Columns are tuples",[790,2048,2049,2050],{},"Multi-row header read as a ",[14,2051,982],{},[790,2053,2054,2055,2057],{},"Flatten with a join, dropping ",[14,2056,1056],{}," parts.",[775,2059,2060,2070,2073],{},[790,2061,2062,2065,2066,2069],{},[14,2063,2064],{},"Region"," column holds a ",[14,2067,2068],{},"Total"," row",[790,2071,2072],{},"Footer not removed",[790,2074,2075,2076,27],{},"Filter by content, not just ",[14,2077,1734],{},[775,2079,2080,2083,2086],{},[790,2081,2082],{},"Works one month, breaks the next",[790,2084,2085],{},"Preamble height changed",[790,2087,2088],{},"Detect the header row by its content.",[775,2090,2091,2094,2099],{},[790,2092,2093],{},"Read is very slow",[790,2095,2096,2098],{},[14,2097,1734],{}," forces a Python parse",[790,2100,2101],{},"Read fully and slice the frame instead.",[775,2103,2104,2109,2112],{},[790,2105,2106],{},[14,2107,2108],{},"ValueError: Passed header=4 but only 3 lines",[790,2110,2111],{},"Sheet shorter than expected, or wrong sheet",[790,2113,2114,2115,2117],{},"Check ",[14,2116,370],{}," and peek first.",[775,2119,2120,2123,2126],{},[790,2121,2122],{},"Numbers read as text",[790,2124,2125],{},"Header row absorbed into the data",[790,2127,2128],{},"Fix the header index; the dtype follows.",[175,2130,2132],{"id":2131},"performance-and-scale-notes","Performance and scale notes",[10,2134,2135,2029,2137,2139,2140,2142,2143,2145],{},[14,2136,657],{},[14,2138,516],{}," do not save any reading. pandas still parses every row of the sheet — the arguments only decide what is kept. The argument that genuinely reduces work is ",[14,2141,1896],{},", which avoids materialising columns entirely, and ",[14,2144,526],{},", which stops early.",[10,2147,2148],{},"For a wide export where you use six of sixty columns, the difference is substantial:",[180,2150,2152],{"className":212,"code":2151,"language":214,"meta":185,"style":185},"import time\nimport pandas as pd\n\nfor label, kwargs in [\n    (\"everything\", {}),\n    (\"six columns\", {\"usecols\": [\"Region\", \"Units\", \"Revenue\",\n                                 \"Owner\", \"Status\", \"Updated\"]}),\n]:\n    start = time.perf_counter()\n    frame = pd.read_excel(\"wide_export.xlsx\", header=4, **kwargs)\n    print(f\"{label:\u003C14} {time.perf_counter() - start:6.2f}s  {frame.shape}\")\n",[14,2153,2154,2161,2171,2175,2186,2197,2225,2243,2248,2258,2283],{"__ignoreMap":185},[189,2155,2156,2158],{"class":191,"line":192},[189,2157,222],{"class":221},[189,2159,2160],{"class":225}," time\n",[189,2162,2163,2165,2167,2169],{"class":191,"line":235},[189,2164,222],{"class":221},[189,2166,226],{"class":225},[189,2168,229],{"class":221},[189,2170,232],{"class":225},[189,2172,2173],{"class":191,"line":242},[189,2174,239],{"emptyLinePlaceholder":238},[189,2176,2177,2179,2182,2184],{"class":191,"line":275},[189,2178,1124],{"class":221},[189,2180,2181],{"class":225}," label, kwargs ",[189,2183,1100],{"class":221},[189,2185,1113],{"class":225},[189,2187,2188,2191,2194],{"class":191,"line":286},[189,2189,2190],{"class":225},"    (",[189,2192,2193],{"class":199},"\"everything\"",[189,2195,2196],{"class":225},", {}),\n",[189,2198,2199,2201,2204,2207,2210,2212,2214,2216,2218,2220,2222],{"class":191,"line":311},[189,2200,2190],{"class":225},[189,2202,2203],{"class":199},"\"six columns\"",[189,2205,2206],{"class":225},", {",[189,2208,2209],{"class":199},"\"usecols\"",[189,2211,292],{"class":225},[189,2213,1690],{"class":199},[189,2215,254],{"class":225},[189,2217,1695],{"class":199},[189,2219,254],{"class":225},[189,2221,1700],{"class":199},[189,2223,2224],{"class":225},",\n",[189,2226,2227,2230,2232,2235,2237,2240],{"class":191,"line":335},[189,2228,2229],{"class":199},"                                 \"Owner\"",[189,2231,254],{"class":225},[189,2233,2234],{"class":199},"\"Status\"",[189,2236,254],{"class":225},[189,2238,2239],{"class":199},"\"Updated\"",[189,2241,2242],{"class":225},"]}),\n",[189,2244,2245],{"class":191,"line":358},[189,2246,2247],{"class":225},"]:\n",[189,2249,2250,2253,2255],{"class":191,"line":364},[189,2251,2252],{"class":225},"    start ",[189,2254,261],{"class":221},[189,2256,2257],{"class":225}," time.perf_counter()\n",[189,2259,2260,2262,2264,2266,2269,2271,2273,2275,2277,2279,2281],{"class":191,"line":400},[189,2261,278],{"class":225},[189,2263,261],{"class":221},[189,2265,509],{"class":225},[189,2267,2268],{"class":199},"\"wide_export.xlsx\"",[189,2270,254],{"class":225},[189,2272,516],{"class":257},[189,2274,261],{"class":221},[189,2276,119],{"class":319},[189,2278,254],{"class":225},[189,2280,1620],{"class":221},[189,2282,1667],{"class":225},[189,2284,2285,2288,2290,2293,2295,2297,2300,2303,2305,2307,2310,2312,2315,2318,2320,2323,2325,2328,2330,2332],{"class":191,"line":405},[189,2286,2287],{"class":319},"    print",[189,2289,637],{"class":225},[189,2291,2292],{"class":221},"f",[189,2294,1577],{"class":199},[189,2296,1550],{"class":1549},[189,2298,2299],{"class":225},"label",[189,2301,2302],{"class":221},":\u003C14",[189,2304,1556],{"class":1549},[189,2306,1396],{"class":1549},[189,2308,2309],{"class":225},"time.perf_counter() ",[189,2311,1811],{"class":221},[189,2313,2314],{"class":225}," start",[189,2316,2317],{"class":221},":6.2f",[189,2319,1556],{"class":1549},[189,2321,2322],{"class":199},"s  ",[189,2324,1550],{"class":1549},[189,2326,2327],{"class":225},"frame.shape",[189,2329,1556],{"class":1549},[189,2331,1577],{"class":199},[189,2333,397],{"class":225},[10,2335,2336,2337,2340],{},"Three habits follow. ",[667,2338,2339],{},"Detect the header once per file, not per sheet"," — the two-pass read costs an extra parse of twenty rows, which is negligible, but running it inside a loop over forty sheets is not. Cache the index when the sheets share a layout.",[10,2342,2343,2349],{},[667,2344,2345,2346,2348],{},"Avoid ",[14,2347,1734],{}," on large files."," It disables the fast path entirely. Filtering by content after a normal read is both faster and more robust.",[10,2351,2352,2355,2356,2360,2361,2365],{},[667,2353,2354],{},"Combine detection with chunked reading for very large sheets."," Find the header from a cheap preview, then stream the body with the approach in ",[23,2357,2359],{"href":2358},"\u002Fadvanced-data-transformation-and-cleaning\u002Fworking-with-large-excel-files-in-python\u002Fread-large-excel-file-in-chunks-with-pandas\u002F","reading large Excel files in chunks with pandas",", so peak memory stays flat regardless of row count. And where the file arrives as a legacy format, convert it first — the parse cost dominates everything above, and ",[23,2362,2364],{"href":2363},"\u002Fgetting-started-with-python-excel-automation\u002Fhandling-excel-file-formats-and-conversions\u002Fconvert-xls-to-xlsx-with-python\u002F","converting .xls to .xlsx"," removes it permanently.",[175,2367,2369],{"id":2368},"conclusion","Conclusion",[10,2371,2372,2373,2375,2376,2378,2379,2381,2382,2384],{},"Reading an Excel export with a title block comes down to knowing that ",[14,2374,759],{}," counts rows in the original file and does the skipping for you, while ",[14,2377,657],{}," renumbers everything after it. Peek at the top with ",[14,2380,582],{}," before writing the real read. Flatten multi-row headers immediately and drop the ",[14,2383,1056],{}," fragments that merged cells leave. And when the preamble height is not stable — which, over enough months, it never is — detect the header row by looking for the column names you expect, and raise a clear error when they are not there.",[175,2386,2388],{"id":2387},"frequently-asked-questions","Frequently asked questions",[10,2390,2391,2399,2401,2402,2404,2405,2408],{},[667,2392,2393,2394,2029,2396,2398],{},"What is the difference between ",[14,2395,657],{},[14,2397,516],{},"?",[14,2400,657],{}," discards rows before pandas looks at the file; ",[14,2403,516],{}," names which of the remaining rows holds the column names. Passing ",[14,2406,2407],{},"header=3"," alone is usually enough, because pandas then treats rows 0 to 2 as ignorable preamble and starts data at row 4.",[10,2410,2411,2414,2415,2417,2418,2420],{},[667,2412,2413],{},"How do I read a file with two stacked header rows?","\nPass a list, for example ",[14,2416,806],{},". pandas builds a ",[14,2419,982],{}," from both rows, which you can then flatten into single names by joining the levels with an underscore.",[10,2422,2423,2426,2427,2429],{},[667,2424,2425],{},"My export has a total row at the bottom — how do I drop it?","\nUse ",[14,2428,1734],{}," with the number of trailing rows to ignore. It requires a Python-level parse, so on very large files it is faster to read everything and slice the frame instead.",[10,2431,2432,2435,2436,2438,2439,27],{},[667,2433,2434],{},"The number of preamble rows changes every month. What then?","\nDo not hard-code it. Read the first twenty rows with ",[14,2437,582],{},", find the row containing your known column names, and pass that index as ",[14,2440,516],{},[10,2442,2443,2450,2451,2453],{},[667,2444,2445,2446,254,2448,2398],{},"Why are my columns named ",[14,2447,16],{},[14,2449,2010],{},"\npandas took a blank row as the header. Either the header index is wrong, or the real header sits below merged title cells. Read with ",[14,2452,582],{}," first and print the top rows to see where the names actually are.",[175,2455,2457],{"id":2456},"related","Related",[2459,2460,2461,2468,2477,2487,2493],"ul",{},[2462,2463,2464,2465,2467],"li",{},"Up to the parent: ",[23,2466,26],{"href":25}," — the full reading toolkit.",[2462,2469,2470,2473,2474,2476],{},[23,2471,2472],{"href":1240},"Handle Merged Cells When Reading Excel with pandas"," — why stacked headers leave ",[14,2475,1056],{}," gaps.",[2462,2478,2479,2483,2484,2486],{},[23,2480,2482],{"href":2481},"\u002Fgetting-started-with-python-excel-automation\u002Freading-excel-files-with-pandas\u002Fread-specific-columns-from-excel-with-pandas\u002F","Read Specific Columns from Excel with pandas"," — the ",[14,2485,1896],{}," argument in depth.",[2462,2488,2489,2492],{},[23,2490,2491],{"href":1723},"Validate Excel Columns Before Import with pandas"," — failing loudly when an export changes shape.",[2462,2494,2495,2498],{},[23,2496,2497],{"href":2358},"Read Large Excel Files in Chunks with pandas"," — combining header detection with streaming.",[2500,2501,2502],"style",{},"html pre.shiki code .sMTad, html code.shiki .sMTad{--shiki-default:#6F42C1;--shiki-dark:#FFB757}html pre.shiki code .srMev, html code.shiki .srMev{--shiki-default:#032F62;--shiki-dark:#ADDCFF}html .default .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html.dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html pre.shiki code .s-kum, html code.shiki .s-kum{--shiki-default:#D73A49;--shiki-dark:#FF9492}html pre.shiki code .skGVy, html code.shiki .skGVy{--shiki-default:#24292E;--shiki-dark:#F0F3F6}html pre.shiki code .sa561, html code.shiki .sa561{--shiki-default:#E36209;--shiki-dark:#FFB757}html pre.shiki code .sP0c6, html code.shiki .sP0c6{--shiki-default:#005CC5;--shiki-dark:#91CBFF}html pre.shiki code .s-wDw, html code.shiki .s-wDw{--shiki-default:#6A737D;--shiki-dark:#BDC4CC}html pre.shiki code .s_Opv, html code.shiki .s_Opv{--shiki-default:#6F42C1;--shiki-dark:#DBB7FF}html pre.shiki code .sSjpA, html code.shiki .sSjpA{--shiki-default:#005CC5;--shiki-dark:#FF9492}",{"title":185,"searchDepth":235,"depth":235,"links":2504},[2505,2506,2507,2508,2509,2510,2511,2512,2513,2514,2515],{"id":177,"depth":235,"text":178},{"id":476,"depth":235,"text":477},{"id":590,"depth":235,"text":591},{"id":849,"depth":235,"text":850},{"id":1244,"depth":235,"text":1245},{"id":1727,"depth":235,"text":1728},{"id":1982,"depth":235,"text":1983},{"id":2131,"depth":235,"text":2132},{"id":2368,"depth":235,"text":2369},{"id":2387,"depth":235,"text":2388},{"id":2456,"depth":235,"text":2457},"2026-08-15","Read Excel files with title blocks, banners and multi-row headers using pandas — skiprows, header, nrows, usecols, and finding the header row automatically when it moves.","md",[2520,2523,2525,2527,2529],{"q":2521,"a":2522},"What is the difference between skiprows and header?","skiprows discards rows before pandas looks at the file; header names which of the remaining rows holds the column names. Passing header=3 alone is usually enough, because pandas then treats rows 0 to 2 as ignorable preamble and starts data at row 4.",{"q":2413,"a":2524},"Pass a list, for example header=[0, 1]. pandas builds a MultiIndex from both rows, which you can then flatten into single names by joining the levels with an underscore.",{"q":2425,"a":2526},"Use skipfooter with the number of trailing rows to ignore. It requires a Python-level parse, so on very large files it is faster to read everything and slice the frame instead.",{"q":2434,"a":2528},"Do not hard-code it. Read the first twenty rows with header=None, find the row containing your known column names, and pass that index as header.",{"q":2530,"a":2531},"Why are my columns named Unnamed 0, Unnamed 1?","pandas took a blank row as the header. Either the header index is wrong, or the real header sits below merged title cells. Read with header=None first and print the top rows to see where the names actually are.",{},"\u002Fgetting-started-with-python-excel-automation\u002Freading-excel-files-with-pandas\u002Fskip-rows-and-set-header-when-reading-excel-with-pandas",{"title":2535,"description":2536},"pandas read_excel: skiprows and header Explained","Handle Excel exports with title rows and stacked headers in pandas — skiprows vs header, callable skiprows, multi-row headers, skipfooter, and auto-detecting the header row.","skip-rows-and-set-header-when-reading-excel-with-pandas","getting-started-with-python-excel-automation\u002Freading-excel-files-with-pandas\u002Fskip-rows-and-set-header-when-reading-excel-with-pandas\u002Findex","how-to","GG0BPplTLa36kl2dgwDk50fgDZUX3bGgCVeBvUJ0ehY",[2542,2546],{"title":2543,"path":2544,"stem":2545,"children":-1},"Read Specific Columns From Excel With Pandas","\u002Fgetting-started-with-python-excel-automation\u002Freading-excel-files-with-pandas\u002Fread-specific-columns-from-excel-with-pandas","getting-started-with-python-excel-automation\u002Freading-excel-files-with-pandas\u002Fread-specific-columns-from-excel-with-pandas\u002Findex",{"title":2547,"path":2548,"stem":2549,"children":-1},"Using openpyxl for Excel File Manipulation","\u002Fgetting-started-with-python-excel-automation\u002Fusing-openpyxl-for-excel-file-manipulation","getting-started-with-python-excel-automation\u002Fusing-openpyxl-for-excel-file-manipulation\u002Findex",1786800027241]