Hi, while the push from CRMA Recipe to AWS S3 Bucket the number of rows changes (e.g. in the dataset in CRMA there're 24mln, the result is 32bln. What is the 'file size', how it is calculated in CRMA? and how to fix the rows multiplication while push to AWS? Thank you!
When exporting data from a Salesforce CRM Analytics (CRMA) recipe to AWS S3, the occurrence of row duplication is usually related to file format settings or data segmentation during export. For example, if CRMA contains 24 million records, but the result in S3 is 32 billion, this may indicate data duplication caused by multiple export attempts, partitioning settings, or specific data formatting in CRMA.
To resolve this issue, first check the export settings in CRMA. Ensure that the data format is compatible with AWS S3 (e.g., CSV or Parquet) and avoid unnecessary segmentation if it's not required. It's also recommended to test the export on a smaller dataset to see if the row multiplication persists.
Additionally, to streamline the export process, you can use third-party solutions such as CloudFiles. These tools provide optimized integration between Salesforce and external data storage, including AWS, helping to avoid data duplication and ensure accuracy in S3.