Part 3 of 9 in After You Call: A Deep Dive on Philly 311
The bigger 311 story isn't the gap. It's the wait.
Lower-income Philadelphia neighborhoods wait 16 days longer for 311. Two-thirds of the gap is the work itself — abandoned cars, blight.
Mean days to close for the 10 largest 311 categories, by income tier. The Y-axis is capped at 160 days to keep routine services legible. Source: OpenDataPhilly 311 Service Requests.
Read the heights of the bars before the gap between them. The three leftmost services — Abandoned Vehicle, Smoke Detector, Traffic Calming — each close in four-and-a-half to five months in both income tiers. The two colors are different, but the difference between them is not the story. The story is that everyone is waiting that long. The 15-day gap on a missed-rubbish pickup is real and worth fixing; it is also the smallest problem on the chart.
In heading into this analysis, I fully expected to find more evidence of disparate impact by income, building on the parent article that found the 2.6 day difference, and I did (in fact, it persists across all major categories per the chart above). But I think I found something more important — that we as Philadelphians are routinely waiting months or years for simple things like towing cars.
A few weeks ago I wrote that Philadelphia 311 requests from lower-income zip codes wait longer to close than the same kinds of requests from higher-income zip codes. The pattern was clear. The mechanism was not. I promised, in that article, to come back and test whether the gap was about the city responding more slowly to the same things, or about the kinds of things each neighborhood was asking for in the first place. This article is that test.
The economics literature calls this the Kitagawa-Oaxaca-Blinder decomposition. It is a way of taking any aggregate gap between two groups — in wages, in test scores, in my case, in mean wait times — and breaking it into two pieces. One piece is the part of the gap that would persist even if the two groups had identical mixes of underlying categories. That is the “rates” component. The other piece is the part of the gap that comes purely from the groups generating different mixes of categories in the first place. That is the “composition” component.
I ran the decomposition on the 311 data. Here is what falls out.
The two possible stories
That gap could come from one of two very different places. The first is what I will call delivery discrimination: a request for a street-light repair in 19142 takes longer to resolve than the identical request in 19103, because crews respond faster to calls from wealthier blocks. The second is what I will call demand composition: lower-income neighborhoods disproportionately generate request types that are inherently slow to resolve (abandoned vehicles, blight remediation, illegal-dumping cleanups), and higher-income neighborhoods disproportionately generate request types that are inherently fast (parking complaints, missed curbside recycling pickups). Same response standard, harder workload.
Both stories can be true. The question is the weight. If delivery discrimination dominates, the city has a service-delivery problem and the right response is some combination of crew reallocation, response-time auditing, and equity mandates. If demand composition dominates, the city is responding at approximately the same per-request rate, and the problem is upstream — the abandoned properties, illegal dumping sites, and deteriorating infrastructure that generate the harder work orders in the first place. Those require a fundamentally different intervention: investment, demolition, and disinvestment remediation.
What the lower-income zips are actually asking for
The next piece of the puzzle is the mix. What share of all 311 requests filed from lower-income zips is for each service type, and how does that share compare to the mix coming from higher-income zips?
Each row is one service category. Orange (left) is the lower-income tier's share; blue (right) is the higher-income tier's share. Rows sorted by gap (lower − higher) so the services where lower-income zips over-index the most sit at the top. The 30+ smaller service categories not shown individually collectively account for 32% of lower-tier requests and 53% of higher-tier requests. Source: OpenDataPhilly 311 Service Requests.
The lower-income zips file Maintenance Complaints at three times the rate of the higher-income zips — 15.6% of all requests vs 5.1%. They file Abandoned Vehicle reports at twice the rate — 7.7% vs 3.5%. They file Smoke Detector complaints at seven times the rate — 1.3% vs 0.2%. The volume differences are not subtle.
Reverse the comparison and the higher-income zips are not filing more of the slow services; they are filing fewer. The categories that dominate the higher-income mix are the routine ones: rubbish and recycling, street defects, traffic. These categories also happen to be the ones with the shortest mean waits.
The decomposition
The Kitagawa-Oaxaca-Blinder decomposition pulls these two threads apart. For each service category, the contribution to the aggregate gap is split into a composition effect (the share difference, weighted by the higher-tier mean) and a rate effect (the mean difference, weighted by the lower-tier share). I add the effects across categories. The math is exact on means: the two effects sum to the observed mean gap with no residual.
The headline number: the observed gap in this restricted analysis is 16.1 days. That gap decomposes to:
The composition effect — the part of the gap that comes from lower-income zips asking for harder things — is roughly 2.5 times larger than the rate effect — the part that comes from the city being slower on identical things. Both effects point the same direction, and both are real, but the demand-mix difference is the dominant force.
Two service categories drive the entire decomposition. Maintenance Complaint and Abandoned Vehicle together account for the bulk of the gap. Maintenance Complaint alone contributes about 9.1 days of the total effect: 6.9 days from composition and 2.2 days from within-service rates. Abandoned Vehicle contributes about 6.8 days: 5.7 from composition and 1.1 from rates. Both are slow, both are heavily skewed toward lower-income zips, and both reflect the physical condition of the built environment.
Top 5 services by contribution to the mean KOB decomposition of the 2024+ 311 wait-time gap, lower-income vs higher-income zips (n=221,803 requests, 40 services). Blue is the composition effect (different mix of requests); amber is the within-service rate effect. Bars are in days. The five services shown account for the bulk of the 16.1-day decomposition; the remaining 35 services contribute fractions of a day each. Source: OpenDataPhilly 311 Service Requests.
The 29% is the interesting part
The 71% gets the headline. The 29% is the more interesting part, because it is the part where the city might be doing something wrong.
Inside the 29%, the largest single per-service rate gap is for Smoke Detector complaints (a 22.5-day within-service gap, though on low volume), followed by Abandoned Vehicle (14.5 days) and Maintenance Complaint (14.3 days). These are large absolute numbers. The reason the rate effect does not dominate the aggregate decomposition is that the slow services are the ones where the within-service rate gap is large — but they are also the ones whose volume is dramatically skewed toward lower-income zips, so the absolute rate-effect contribution stays small.
Some of the within-service rate gaps point the other way. A pothole report in a lower-income zip closes in 50.9 days on average; the same report in a higher-income zip closes in 60.1 days. That is a 9.2-day gap, but in the higher-income direction. The mean is dominated by the long tail of bad cases — the multi-year potholes, the abandoned work orders, the cases that get stuck in the system — and on those cases, the higher-income zips do worse. The pattern reverses entirely on Street Light Outage too (lower tier is faster, by a wide margin). The rate effect is not uniform, and it is not safe to read it as “the city is slower to poor neighborhoods on the same work.” It is messier than that.
I want to be careful with what this means. The decomposition is a useful frame, but it is a frame. There is a third possibility that the math cannot distinguish: the city's response standard is the same, but the work is harder in lower-income zips. An abandoned vehicle on a block with no off-street parking, where the owner cannot be located, where the title is in probate — that takes longer to resolve regardless of which neighborhood it is in. The “rate” gap I am measuring on the slow services may not be a delivery-discrimination gap. It may be a residual difficulty gap, after controlling for the category.
If the 29% is mostly a difficulty gap rather than a discrimination gap, then the policy response is not crew reallocation. It is something harder. The city would need to invest in the tools, staffing, and process re-engineering that makes hard work less hard — or, more bluntly, accept that some of these requests will take a long time no matter who is responding. A smoke-detector complaint in a building with no working smoke detectors in any unit is not the same kind of work as one in a building where the detector just needs a new battery. The 311 system cannot tell the difference, and the mean reflects that.
What the city should and should not do
The decomposition tells me three things.
Don't reallocate crews — the gap isn't on routine services.
The city does not have a 311 discrimination problem on routine services. Rubbish collection, street-light repair, street defects, illegal dumping — these close at near-equal speeds in both kinds of neighborhoods. Reallocating crews based on the parent article's headline would be solving the wrong problem.
Fix the built environment, not response times — this is a capital-projects problem.
The city cannot fix the demand mix by rebalancing crews. It can fix the demand mix by fixing the conditions that generate the demand — demolishing the abandoned properties, towing the abandoned vehicles, repairing the alleys and sidewalks that flag as maintenance complaints. Those interventions are capital projects, not service-level operations. They live in the capital budget, the demolition pipeline, the land bank. They are slower, more expensive, and politically harder than tweaking 311 response standards.
Equal response times won't fix the 150-day wait.
Even if the city equalized response times across neighborhoods, the absolute problem would remain. The abandoned vehicle in 19142 has been there for 150 days on average. The one in 19103 has been there for 135. Both are problems. Neither is solved by rebalancing crews.
When I lived in Strawberry Mansion, I watched abandoned cars sit on the block for what felt like the better part of a year. The 15-day gap between my zip and Society Hill's was real and worth fixing. But it was not what I noticed, and I don't think it would have been what I cared about even if I had. What I noticed was that the car was still there.
Most people, rich or poor, feel the absolute wait, not the relative one.
The standard the city should be measuring is not the response time. The standard is what the city is being asked to do — and that varies by neighborhood in ways the 311 system cannot fix on its own.
Notes, Sources, and Methodology
All data is the OpenDataPhilly 311 Service Requests table public_cases_fc, pulled 2026-06-08 and restricted to filings from 2024-01-01 forward, excluding service_name = 'Information Request' and any ticket where the elapsed days between filing and closure fell outside [0, 365]. The cohort used in this article is 221,803 closed requests from 24 zip codes across the lower- and higher-income tiers used throughout the series (lower < $35K, higher ≥ $65K median household income, ACS 2022 5-year estimates).
Methodology. The Kitagawa-Oaxaca-Blinder decomposition is the standard decomposition of a group-mean gap (Blinder 1973, "Wage Discrimination"; Oaxaca 1973, "Male-Female Wage Differentials"). KOB on means is exact: the composition and rate effects sum to the observed gap with no residual. I use the focal-group-share weighting for the rate effect. Service categories with fewer than 30 requests in either tier are excluded from the decomposition, leaving 40 of the 49 categories that appeared in the cohort.
Limitations. The middle-income tier is held out of this article. The “All other” service categories (the 9 of 49 excluded for thinness) include small-volume codes whose means are noisy; their aggregate contribution to the gap is fractions of a day and does not change the 71/29 split. The decomposition answers the question it sets out to answer — what share of the gap is the mix vs. the response rate — but it does not separate true delivery discrimination from residual case difficulty (a harder abandonment case takes longer no matter who responds), which the public data cannot distinguish.