{
  "contents": [
    {
      "role": "user",
      "parts": [
        {
          "text": "Attachment 1 is a complete English sung advertorial story-song; the other attachment(s) are complete Ukrainian candidates telling the same story with a different product. Listen to every whole song end to end. Compare at the META level with at least 10 timed landmarks per attachment: (a) word attacks and releases (crisp consonant endings vs sustained open vowels; which words are lengthened and whether they are the important ones); (b) register: are asides and connectives dropped to a speech-like chest register while the top is saved for a few words; (c) melodic play with words (a question flip, a surprise, a landing); (d) breath and phrase endings; (e) how the arrangement reacts to story turns (a stop on the advice word, thinning on the time word, air before the brand); (f) whether Larysa, the daughter and the ex-husband's wife sound like different people in one voice; (g) pacing of information (fast lists, slow payoffs, the product/order block); (h) IMPORTANT: is each candidate genuinely SUNG throughout (melodic, pitched, natural Ukrainian singing) or does it slip into spoken recitation / rap-like talking, and where exactly (local seconds)? A low chest register is desired for asides, but the owner requires real singing. Then give per candidate: sungNotRecited{verdict,recitedPassages[{attachment,localSeconds,words}]}, closerToSourceThanTypicalUkrainianPop[], remainingGaps[], defects[{attachment,localSeconds,words,problem,severity}], wordsMisheardOrUnclear[], brandNameHeard, allOrderConditionsHeard(boolean), verdictClosestToSource(attachment number) with why. Return ONLY JSON (no prose outside JSON), <=1900 words, English. Required keys: mediaAccess(boolean), audioAccess(boolean), attachmentsHeard[{attachment,firstWordsHeard,lastWordsHeard,durationEstimateSeconds}], then the analysis keys listed below. No script is supplied: quote only words you actually hear and mark uncertain words with (?). Use attachment number plus LOCAL seconds of that attachment. Do not rate quality with numbers, do not assume either language or the original is better, do not reward louder mastering, more notes or softer timbre. Identify concrete mechanisms, not taste."
        },
        {
          "text": "Attachment 1: Attachment 1; complete English original song, 194.5s, 56kbps mono proxy with fixed gain; local time starts at 0."
        },
        {
          "inlineData": {
            "mimeType": "audio/mpeg",
            "data": "[base64 of frozen bytes; see inputs.json]"
          }
        },
        {
          "text": "Attachment 2: Attachment 2; complete Ukrainian candidate A, 186.3s, 56kbps mono proxy with fixed gain; local time starts at 0."
        },
        {
          "inlineData": {
            "mimeType": "audio/mpeg",
            "data": "[base64 of frozen bytes; see inputs.json]"
          }
        },
        {
          "text": "Attachment 3: Attachment 3; complete Ukrainian candidate B, 203.2s, 56kbps mono proxy with fixed gain; local time starts at 0."
        },
        {
          "inlineData": {
            "mimeType": "audio/mpeg",
            "data": "[base64 of frozen bytes; see inputs.json]"
          }
        }
      ]
    }
  ],
  "stream": true,
  "generationConfig": {
    "thinkingConfig": {
      "includeThoughts": false,
      "thinkingLevel": "high"
    }
  }
}
