jastorj/couchmind-v5.8.2_cold_start-cw-26K-16bit
0518
Snowflake/Arctic-Text2SQL-R1-7B Fine-tuned for NL2SQL++ v5.8.2coldstart (SFT)
This model is a supervised fine-tuned version of Snowflake/Arctic-Text2SQL-R1-7B on the NL2SQL++ v5.8.2coldstart dataset with code-with-thought reasoning.
Model Details
- Base Model: Snowflake/Arctic-Text2SQL-R1-7B
- Task: Text-to-SQL generation
- Dataset: NL2SQL++ v5.8.2coldstart with code-with-thought reasoning
- Fine-tuning Method: LoRA (Low-Rank Adaptation) with Unsloth
- Quantization: 16-bit merged weights
- Training Dataset Size: 1148 examples
- Validation Dataset Size: 0 examples
Full training configuration: see training_args.json in this repo.
Train Dataset Example
<|im_start|>system
You are a helpful AI Assistant that provides well-reasoned and detailed responses. You first
think about the reasoning process as an internal monologue and then provide the user with
the answer. Respond in the following format: <think>
...
</think>
<answer>
...
</answer><|im_end|>
<|im_start|>user
You are a Couchbase SQL++ query expert. Given a database schema and a natural language question, generate a syntactically valid SQL++ query that precisely answers the question.
Rules:
- SELECT only the columns explicitly asked for — nothing more, nothing less
- Use the exact bucket, scope, and collection names provided in the database schema
Bucket Name:
`bird_training_bucket`
Scope Name:
`beer_factory`
Database Schema:
{"`bird_training_bucket`.`beer_factory`.`geolocation`": {"properties": {"LocationID": {"samples": [1, 0], "type": "number"}, "Longitude": {"samples": [-121.42243, -121.426432], "type": "number"}, "Latitude": {"samples": [38.559615, 0], "type": "number"}}, "type": "object"}, "`bird_training_bucket`.`beer_factory`.`rootbeerbrand`": {"properties": {"Twitter": {"samples": ["https://twitter.com/ilovedads", "https://twitter.com/1919rootbe..."], "type": ["null", "string"]}, "Description": {"samples": [" Made In Maine Sweetened with ...", "AJ Stephans Company makes \nthe..."], "type": "string"}, "City": {"samples": ["Bundaberg", "Bainbridge Island"], "type": ["null", "string"]}, "BreweryName": {"samples": ["Abita Brewery", "AJ Stephans Beverages"], "type": "string"}, "AvailableInKegs": {"samples": ["TRUE", "FALSE"], "type": "string"}, "BrandName": {"samples": ["A.J. Stephans", "Abita"], "type": "string"}, "BrandID": {"samples": [10002, 10003], "type": "number"}, "Country": {"samples": ["Australia", "United States"], "type": "string"}, "CornSyrup": {"samples": ["FALSE", "TRUE"], "type": "string"}, "FacebookPage": {"samples": ["https://www.facebook.com/Gales...", "https://www.facebook.com/Dog-N..."], "type": ["null", "string"]}, "CaneSugar": {"samples": ["FALSE", "TRUE"], "type": "string"}, "Caffeinated": {"samples": ["FALSE", "TRUE"], "type": "string"}, "Website": {"samples": ["http://www.ajstephans.com/", "http://reedsinc.com/product/vi..."], "type": ["null", "string"]}, "AvailableInCans": {"samples": ["FALSE", "TRUE"], "type": "string"}, "FirstBrewedYear": {"samples": [1919, 1898], "type": "number"}, "Alcoholic": {"samples": ["FALSE"], "type": "string"}, "ArtificialSweetener": {"samples": ["TRUE", "FALSE"], "type": "string"}, "CurrentRetailPrice": {"samples": [3, 1], "type": "number"}, "State": {"samples": ["DE", "CA"], "type": ["null", "string"]}, "AvailableInBottles": {"samples": ["FALSE", "TRUE"], "type": "string"}, "Honey": {"samples": ["FALSE", "TRUE"], "type": "string"}, "WholesaleCost": {"samples": [0.4, 0.98], "type": "number"}}, "type": "object"}, "`bird_training_bucket`.`beer_factory`.`rootbeer`": {"properties": {"ContainerType": {"samples": ["Can", "Bottle"], "type": "string"}, "PurchaseDate": {"samples": ["2014-12-06", "2015-03-18"], "type": "string"}, "BrandID": {"samples": [10003, 10002], "type": "number"}, "LocationID": {"samples": [2, 1], "type": "number"}, "RootBeerID": {"samples": [100477, 100389], "type": "number"}}, "type": "object"}, "`bird_training_bucket`.`beer_factory`.`location`": {"properties": {"StreetAddress": {"samples": ["6000 J St"], "type": ["null", "string"]}, "City": {"samples": ["Sacramento"], "type": ["null", "string"]}, "LocationID": {"samples": [0, 1], "type": "number"}, "LocationName": {"samples": ["LOST", "Sac State American River Court..."], "type": "string"}, "State": {"samples": ["CA"], "type": ["null", "string"]}, "ZipCode": {"samples": [95819], "type": ["null", "number"]}}, "type": "object"}, "`bird_training_bucket`.`beer_factory`.`transaction`": {"properties": {"RootBeerID": {"samples": [100229, 100246], "type": "number"}, "TransactionID": {"samples": [101906, 101487], "type": "number"}, "CreditCardType": {"samples": ["Discover", "American Express"], "type": "string"}, "PurchasePrice": {"samples": [1, 3], "type": "number"}, "CreditCardNumber": {"samples": [4916435869897994, 371009703397242], "type": "number"}, "CustomerID": {"samples": [143878, 194680], "type": "number"}, "TransactionDate": {"samples": ["2015-05-26", "2015-07-02"], "type": "string"}, "LocationID": {"samples": [1, 2], "type": "number"}}, "type": "object"}, "`bird_training_bucket`.`beer_factory`.`customers`": {"properties": {"Email": {"samples": ["b.john@outlook.com", "a_ryan19@mail.com"], "type": "string"}, "Last": {"samples": ["Anderson", "Bowker"], "type": "string"}, "CustomerID": {"samples": [230467, 214537], "type": "number"}, "PhoneNumber": {"samples": ["(916) 216-5687", "(916) 184-8050"], "type": "string"}, "State": {"samples": ["CA"], "type": "string"}, "First": {"samples": ["Anna", "Cecil"], "type": "string"}, "SubscribedToEmailList": {"samples": ["FALSE", "TRUE"], "type": "string"}, "City": {"samples": ["Antelope", "Carmichael"], "type": "string"}, "FirstPurchaseDate": {"samples": ["2012-09-24", "2012-10-21"], "type": "string"}, "StreetAddress": {"samples": ["224 Tennis Ct", "114 Parkside Ct"], "type": "string"}, "ZipCode": {"samples": [94279, 94234], "type": "number"}, "Gender": {"samples": ["F", "M"], "type": "string"}}, "type": "object"}, "`bird_training_bucket`.`beer_factory`.`rootbeerreview`": {"properties": {"BrandID": {"samples": [10003, 10002], "type": "number"}, "CustomerID": {"samples": [161175, 105771], "type": "number"}, "ReviewDate": {"samples": ["2013-01-30", "2012-09-12"], "type": "string"}, "Review": {"samples": ["Cans? Not a chance.", "Awww, doggies!"], "type": ["null", "string"]}, "StarRating": {"samples": [2, 1], "type": "number"}}, "type": "object"}}
This schema describes the structure of the data in the specified bucket and scope. It includes information about the collections, fields, and their data types.
Question:
review content of "Tastes like Australia." refers to Review = 'Tastes like Australia.';
For the customer who leaves the review content of "Tastes like Australia.", when was his/her first purchase date?<|im_end|>
<|im_start|>assistant
<think>
I start by locating where the review text lives. The question explicitly talks about review content, and in the schema the collection that has a Review field is `rootbeerreview`. That collection also has `CustomerID`, which is important because I need to connect the review back to the person who wrote it.
Then I look for where “first purchase date” is stored. There are two possible ways to read that:
- I could compute a customer’s earliest purchase from the `transaction` collection using `TransactionDate`.
- Or I could use the `FirstPurchaseDate` field that already exists in the `customers` collection.
Since the schema gives me a direct `FirstPurchaseDate` column in `customers`, that is the cleaner and more literal match to “when was his/her first purchase date?” So I do not need an aggregate like MIN over transactions here.
So the two relevant collections are:
- `customers`: has `CustomerID` and `FirstPurchaseDate`
- `rootbeerreview`: has `CustomerID` and `Review`
That means I need a join on `CustomerID`. I do not need `rootbeer`, `rootbeerbrand`, `location`, or any of the other collections, because the question is only about identifying the customer by a review text and then reading that customer’s first purchase date.
Next, I need the exact filter value for the review. The question says review content of “Tastes like Australia.” and even clarifies that this refers to `Review = 'Tastes like Australia.'`. So the comparison should be against that exact string, including the period at the end.
Now for the output: the question asks “when was his/her first purchase date?” so I only need the `FirstPurchaseDate` field from the customer side of the join.
I’ll alias `customers` as `T1` and `rootbeerreview` as `T2` just to keep the field references short and clear. Then the structure is:
- select `T1.FirstPurchaseDate`
- from `customers` as `T1`
- inner join `rootbeerreview` as `T2`
- on `T1.CustomerID = T2.CustomerID`
- where `T2.Review = 'Tastes like Australia.'`
I briefly consider whether I need `DISTINCT`, because in principle more than one review row could have the same text. But the wording says “the customer” in the singular, and there is nothing else in the question asking me to deduplicate or aggregate, so a straightforward inner join with the review-text filter is enough.
So I arrive at selecting the customer’s `FirstPurchaseDate` from `customers`, joining to `rootbeerreview` by `CustomerID`, and filtering the review text to exactly `Tastes like Australia.`
</think>
<answer>
SELECT T1.`FirstPurchaseDate` FROM `bird_training_bucket`.`beer_factory`.`customers` AS T1 INNER JOIN `bird_training_bucket`.`beer_factory`.`rootbeerreview` AS T2 ON T1.`CustomerID` = T2.`CustomerID` WHERE T2.`Review` = 'Tastes like Australia.'
</answer><|im_end|>
