I recently launched SignalCards Evidence, a free Shopify app focused on one specific type of chargeback: item not received.
I built it after looking at how merchants prepare chargeback responses. One thing I noticed is that you can have a lot of information about an order, but still not know clearly what evidence is actually available, what still needs to be confirmed, and what is missing.
The app lets you import a Shopify order or create a case manually. It then helps organize things like:
order, payment, fulfillment and refund information
available, unconfirmed and missing evidence
things the merchant may still need to check
the case timeline
a final human review before preparing a printable response
It doesn’t submit the dispute automatically and it doesn’t promise that a chargeback will be won. The merchant still reviews everything and submits the response.
The app is currently free on the Shopify App Store:
I’m especially interested in hearing from merchants who have already dealt with an item-not-received chargeback, or who currently have one open.
At this stage, practical feedback would be much more useful to me than reviews. For example, is there information that’s difficult to collect? Is anything in the workflow unnecessary? Is there something that would make preparing the response faster?
Thanks, and I’m happy to answer questions about how it works.
Thanks, this is exactly the kind of feedback I was hoping to get.
The completed-case point is especially useful. I agree that “the checklist feels clear” is not enough. I need to see whether the app actually reduces preparation time, surfaces missing evidence earlier, and produces something merchants are willing to use in a real response.
Your point about evidence outside Shopify is important too. Carrier details, signatures, customer messages, address changes and refund timing are exactly where the case can become fragmented. I also like the idea of being much stricter about provenance: what came from Shopify, what came from the merchant, and what is still only an unconfirmed assertion.
I’m going to start tracking the case-level metrics you suggested rather than just collecting qualitative feedback.
If you’ve handled an item-not-received case yourself, even a completed one, I’d be very interested in having you run it through SignalCards Evidence and tell me where the workflow breaks down or where the source/provenance is unclear. That would be more useful to me right now than a review.
I cannot take your ask, because I have not handled an item not received case myself. So this is about the line where you said you are moving to case level metrics rather than qualitative feedback, since that decision is about to shape everything you learn next.
Two of the three you named will not measure what you want. Whether the app REDUCES PREPARATION TIME needs a without-the-app number, and you will never get one from the same merchant on the same case, because they only prepare it once. Whether it SURFACES MISSING EVIDENCE EARLIER has the same problem: earlier than what, when the counterfactual never happened. Both will produce numbers, and the numbers will mostly measure who chose to use it.
The third one is self contained and worth more than the other two together. Of the items the app marked MISSING, how many did the merchant actually go and obtain before submitting. That needs no control group and no baseline, and it tests the thing that matters, which is whether the list is ACTIONABLE rather than merely accurate. A checklist that is completely right and moves nobody is the failure mode you cannot see from qualitative feedback.
And the case worth chasing hardest is the one where the app showed everything available and the response still lost. One of those tells you more than ten that felt clear.
That’s a fair correction. I was treating preparation time and earlier gap detection as outcome metrics when, on a single case, there’s no real counterfactual to compare them against. I can still collect preparation time as descriptive context, but I shouldn’t interpret it as time saved.
The actionable-missing-evidence measure is much cleaner. If the app flags something as missing, did the merchant actually obtain it before submitting? That directly tells me whether the gap detection changes the case rather than just describing it.
And I agree on the “everything available, still lost” cases. Those are probably the ones most likely to expose something the model or workflow is missing.
I’m going to adjust what I track around those two things. Thanks for catching that before I built the validation around the wrong numbers.
One correction to my own suggestion, because the metric I handed you has a smaller version of the same problem I was complaining about.
Of the items flagged missing, how many the merchant obtained.
A merchant who goes and gets a signature confirmation may have been going to get it anyway, and you cannot tell from the count. It is a weaker version of the counterfactual I said you could not have.
There is a control available inside the same case though, which is why I still think it is the right measure. On any given case some flagged items get obtained and some do not. So the comparison is not app versus no app, it is obtained versus not obtained within one case, against the outcome.
If the cases that lost are disproportionately the ones where a flagged item stayed unobtained, the flag is doing work.
If losses fall evenly across both, the list is accurate and inert, which is the thing you cannot see from anybody telling you the checklist feels clear.
On the everything available and still lost cases.
That set is only readable if you have captured which reason code the dispute was filed under, because the required evidence is defined per code rather than per dispute type.
Item not received is not one thing. Everything available is a sentence about your data.
Everything required is a sentence about the network rules, and only the second one predicts an outcome.
So I would record the reason code on every case from now, even before you know what you will do with it. It is the field you cannot backfill.
Thanks, this is helpful. The distinction between triaging a backlog and building a reference makes sense.
I’ll keep the networks separate, use each network’s own category structure, and treat the code as the index rather than the content.
I also agree on scope. Three or four codes done properly from the primary documents, with the edition cited, sounds much more useful than trying to build a complete table too early.